Skip to main content
If your model server speaks an OpenAI-compatible API, the Barndoor LLM Gateway can route to it. That covers vLLM, Ollama, SGLang, LM Studio, llama.cpp’s llama-server, and in-house wrappers alike. Connecting one is configuration on both sides — no adapter, no SDK, no code. What you get for doing it: every request to your own GPUs carries the same user identity, budget enforcement, rate limits, and audit record as a request to OpenAI or Anthropic. Your developers point at one endpoint and stop caring which side of the line a model lives on. And if your self-hosted capacity runs out, a model route can fall back to a hosted provider automatically — see Failover, Cooldowns, and Route Health.
This guide assumes you already have an admin account and know your way around LLM Management. If not, read Using the LLM Gateway first.
Only Step 1 differs between servers — everything after it is identical, because Barndoor sees the same OpenAI-compatible endpoint either way.

Before you start

You’ll need three things:
  • A running model server, reachable over the network from Barndoor (see Making the server reachable — this is the step that trips people up).
  • The API key you started that server with.
  • Admin access to LLM Management in the Barndoor portal.

Step 1: Start your server so Barndoor can reach it

Two defaults bite on nearly every model server, and both must change before Barndoor can use it:
  1. It listens on loopback only, so nothing off that machine can connect.
  2. It requires no credential, so anything that can connect may use your GPUs freely.
--host defaults to 127.0.0.1. Without --host 0.0.0.0, vLLM binds to loopback only and every connection from outside that machine — including Barndoor’s — is refused. This is the single most common cause of a provider that saves as Enabled · not serving.
--api-key turns on bearer-token authentication for the /v1 paths. Treat the value as a secret: it is the only thing standing between your GPUs and anyone who can reach the port. Barndoor stores it encrypted and never exposes it to end users.Flags worth knowing about:
By default vLLM advertises the model under its Hugging Face repo id, so vllm serve Qwen/Qwen3-8B reports Qwen/Qwen3-8B in /v1/models and expects that exact string in the model field of a request.--served-model-name qwen3-8b overrides that with something friendlier. Whatever you choose, it must match what you enable in Barndoor in Step 3. The gateway passes the name through untouched.A name without a / is also easier to call directly. The gateway reads everything before the first / in a request’s model field as a provider name, so a model named Qwen/Qwen3-8B can only be called through a Model Route (Step 4) or as <provider name>/Qwen/Qwen3-8B.
vLLM only emits OpenAI-style tool_calls when the server was started for it, and the correct parser depends on the model family:
Without these flags the model will describe a tool call in prose instead of returning a structured one, and clients that expect tool use will appear to hang or loop. Barndoor passes tools through unchanged either way — it can’t compensate for a server that wasn’t started for tool calling.
For reasoning models, --reasoning-parser qwen3 (or deepseek_r1, depending on the family) splits thinking tokens into their own field rather than leaving them inline in content. Recommended if your clients render responses directly to users.
Sets the maximum combined prompt and output length. If unspecified, vLLM derives it from the model config, which can be more than your GPUs can actually hold. Setting it explicitly turns an out-of-memory crash into a clean HTTP error that Barndoor reports and can fail over from.
Confirm the server answers before you go anywhere near the portal:
You should get a JSON list containing the model id you intend to use. Run this from a different machine than the one hosting the server — running it locally passes even when the loopback-binding problem above is present, which is exactly how that problem stays hidden.

Making the server reachable

Barndoor connects to your server over the network like any other client, which means the gateway has to be able to route to it. This is independent of which server you run.
A server on a private subnet is not reachable from app.barndoor.ai. You need to give it a routable address:
  • Put it behind your own load balancer or ingress with a public DNS name and TLS.
  • Restrict who can reach that address — an IP allowlist for Barndoor’s egress addresses, or mutual TLS at your edge. --api-key is authentication, not network isolation.
The certificate must be issued by a publicly trusted CA. The gateway validates TLS and will reject a self-signed or private-CA certificate — it surfaces as could not reach upstream provider, which reads like a connectivity problem rather than a certificate one.

Step 2: Add the provider

In the Barndoor portal, go to LLM Management → Providers, click Add Provider, and on the Custom Provider card click Add Custom. Don’t pick one of the named vendor cards. That’s the right choice rather than a compromise. The named cards carry facts about a vendor’s hosted endpoint: its URL, the models it serves, its per-token prices. None of those are knowable for a server you run yourself. The Add Custom Provider form asks for the things that actually vary and assumes nothing else, which is why it works the same for every server in Step 1. The base URL and key belong to the credentials, not the provider. To rotate the key or change the address later, edit the credentials on the Credentials tab. Every provider using them picks up the change, and a changed address triggers a fresh health check.
The base URL must not end in /v1.vLLM’s own documentation shows http://localhost:8000/v1 because that’s what OpenAI SDK clients expect. Barndoor is not an SDK client. It appends the version segment itself, requesting {base_url}/v1/models and {base_url}/v1/chat/completions, so a /v1 on the end would produce /v1/v1/… on every request. Barndoor refuses to save it: the form shows base_url must not end in '/v1' along with the corrected URL to use instead.

How Barndoor checks the connection

On save, Barndoor runs a connectivity check: a GET {base_url}/v1/models with your key, which times out after 10 seconds. The result decides the status shown under the provider’s name in the Providers list: A provider that is not serving stays suspended from routing until a later check passes. Checks run when you save the provider and when you choose Re-check health from the provider’s row menu (⋮). Turn Block traffic on failed health check off only if your server sits behind something that blocks the model-list endpoint but serves completions fine.

Step 3: Enable the models

Right after you save the provider, the Enable Models dialog opens. A custom provider has no catalog to pick from, which is the honest state of affairs: only your server knows what it loaded. Type your model name into Custom Models (optional). You can enter several names, separated by commas or new lines. If you skip this step, open the provider later and click Add Models. The name must match what /v1/models reports, exactly. Take it from the server rather than from memory:
That’s the Hugging Face repo id (Qwen/Qwen3-8B) unless you set --served-model-name, in which case it’s whatever you chose. Copy it character for character — Qwen/Qwen3-8B and qwen3-8b are different models as far as the gateway is concerned, and a mismatch surfaces only when the first request 404s. Models added this way carry a Custom badge. That’s provenance, not a warning: you typed the name rather than picking it from the Barndoor catalog. It’s also the first place to look when requests to a model fail, since a typo looks exactly like a real model until the first request.
If your organization requires pricing before a model can be enabled, saving a self-hosted model fails with Model '…' has no pricing. Either add a rate under Model Pricing first, or turn off Calculate token cost on the provider. See Cost reporting.

Step 4: Create a model route

A model you just enabled is callable only in <provider name>/<model> form, which ties clients to this one provider. Under LLM Management → Model Routes, click Create Route, enter the Route Name your developers will actually use (say qwen3), and point it at the model you just enabled. Developers then call it as qwen3. This indirection is what lets you move traffic later without touching a single client: repoint the route at a bigger GPU node, or add a hosted provider as a second target so requests spill over when your own capacity is saturated.

Step 5: Verify the connection

Test in layers. Each one isolates a different failure, and a green result at one layer tells you nothing about the next — so resist skipping ahead when something breaks. Almost every failed setup is diagnosed by finding the lowest layer that fails.
1

Your server is serving — on the server host

Proves the process is up and the weights finished loading. A large model can take minutes; until it’s ready this returns nothing useful no matter how correct your configuration is.
2

It's reachable from somewhere else

Run the same curl from a different machine — your laptop, a bastion, anything that isn’t the server host — against the exact address you plan to give Barndoor:
Do not skip this by testing on the server host. Loopback succeeds even when the server is bound to 127.0.0.1 and unreachable by everything else, which is the most common reason a provider saves as Enabled · not serving. This layer is the only one that catches it, along with firewall and security-group problems.
3

Barndoor can reach it

Saving the provider runs the connectivity check (see How Barndoor checks the connection). To re-test after changing something on your side, open the provider’s row menu (⋮) in LLM Management → Providers and choose Re-check health. The result appears as a notification, for example Health check passed or Health check failed: …. Health is only re-evaluated on save and on Re-check health, so a server you just fixed keeps its old status until you do one of them.
4

The model is exposed to a caller

Create a key under Settings → My Models → API Keys, then ask the gateway what that key can actually use:
Your route name should be listed, with your provider’s name in its barndoor.provider field. This is a different question from the previous layer: it proves the route resolves and that model access policy lets this caller use it.If the provider is healthy but the model isn’t listed here, the problem is the route or the access policy, not your server. Go back to Step 4, and check LLM Management → Model Access.
5

A request completes — both ways

Then run it again with "stream": true added. Streaming goes through a different path, an SSE relay with its own idle timeout, and it’s what coding assistants and chat UIs actually use. A passing non-streaming request is not evidence that streaming works.
Expect incremental data: chunks, a final chunk carrying usage, then data: [DONE]. The usage chunk is what makes a streamed request countable. The gateway asks for it on every streamed request even when the client doesn’t, so your server must accept stream_options. If the usage chunk is missing, tokens won’t reach your reports.
6

Governance actually applied

This is the layer that justifies the gateway existing, and the one most people forget to check.Open the LLM Usage Dashboard and confirm your requests appear, attributed to the user whose key sent them, with a token count. Cost will read zero, which is expected for self-hosted until you set rates (see below). See Reading the LLM Usage Dashboard.For real assurance that controls apply to self-hosted traffic the same as hosted, set a deliberately tiny token budget scoped to your test key under LLM Management → Budgets, with action Block, and confirm the next request is refused. Budget definitions are cached for up to five minutes, so allow for that before concluding it doesn’t work. See LLM Controls.

Cost reporting for self-hosted models

Self-hosted inference has no market rate (the cost is your own GPU time), so Barndoor ships no default pricing for self-hosted models. Until you decide, usage reporting counts tokens accurately, reports the cost as zero, and shows the model as Unpriced. You have two ways to make that a decision rather than a gap:
  • Record self-hosted traffic as not metered. Edit the provider, turn off Calculate token cost, and set Billing arrangement to Self-hosted. Tokens are still counted, cost is recorded as $0 on purpose, and the model shows Not metered instead of Unpriced. The Billing arrangement is required when cost calculation is off.
  • Give it a rate. Leave Calculate token cost on and set your own rates under LLM Management → Model Pricing. A reasonable approach is to divide the fully loaded hourly cost of the instance by the tokens it produces in that hour, and enter the result as an input and output rate. It doesn’t need to be exact to be useful: even a rough number makes “what did this team actually consume” answerable, and lets a cost budget act as a real ceiling. See Managing Model Pricing.
Token budgets and rate limits apply either way.

Troubleshooting

Barndoor couldn’t open a connection. In rough order of likelihood:
  • vLLM is bound to loopback — restart it with --host 0.0.0.0.
  • A firewall or security group blocks the port from Barndoor’s side.
  • You’re on Barndoor SaaS and the address is private. See Making the server reachable.
  • You’re using https:// with a self-signed or private-CA certificate.
Test from a third machine, not the server host — that’s what distinguishes “not listening publicly” from “not running”.
Something is accepting the connection but not answering in time. Usually a load balancer or proxy in front of your server that’s routing to a dead backend, or a server still loading model weights. A large model can take minutes to become ready. Wait for the server’s own logs to report it’s serving, then choose Re-check health on the provider.
The key in Barndoor doesn’t match the token your server was started with. Check for a trailing newline or a shell-quoting artifact: these are usually copy-paste damage rather than the wrong secret. To replace the key, edit the credentials on the Credentials tab. Barndoor also raises a critical admin alert when a health check finds a provider’s credential rejected.
The save was refused because the base URL ends in /v1. Use the corrected URL the message suggests, which is the same address without the trailing /v1. The gateway adds /v1 itself.
The health check treats a missing model list (HTTP 404 or 405) as inconclusive, so the provider still shows Enabled. If requests then 404 too, check whether a reverse proxy in front of your server is rewriting or blocking /v1 paths, and whether the base URL points at the host root rather than a sub-path.
The name enabled in Barndoor doesn’t match what the server serves. Compare against curl http://<host>:8000/v1/models: the id there is the name to enable, letter for letter. If you added --served-model-name after configuring Barndoor, the enabled name is now stale.If the gateway itself answers 404 model '…' not found, the problem is the name the client sent, not the server. A model ID containing / (such as Qwen/Qwen3-8B) can’t be called by its bare name. Call it through a Model Route, or as <provider name>/Qwen/Qwen3-8B.
The server wasn’t started with --enable-auto-tool-choice and a --tool-call-parser matching the model family. Barndoor forwards the tools array unchanged; it cannot synthesize structured tool calls from a server that isn’t producing them.

Current limits

  • Chat completions and streaming are the supported surface. Claude Code and other Anthropic-format clients can use a self-hosted model too, because the gateway translates Messages requests into chat completions. The gateway also forwards embeddings, legacy completions, and Responses API requests to your server unchanged, but those aren’t verified on this path and work only if your server implements them.
  • Tool calling and reasoning output depend entirely on how you started the server, so Barndoor makes no promise about them on your behalf. They pass through when your server produces them.
  • Barndoor SaaS cannot reach a server on a private network. There is no tunnel or agent for this today; the server needs a routable address, or Barndoor needs to run inside your infrastructure.