llama-server, and in-house wrappers alike. Connecting one is configuration on both sides — no adapter, no SDK, no code.
What you get for doing it: every request to your own GPUs carries the same user identity, budget enforcement, rate limits, and audit record as a request to OpenAI or Anthropic. Your developers point at one endpoint and stop caring which side of the line a model lives on. And if your self-hosted capacity runs out, a model route can fall back to a hosted provider automatically — see Failover, Cooldowns, and Route Health.
Before you start
You’ll need three things:- A running model server, reachable over the network from Barndoor (see Making the server reachable — this is the step that trips people up).
- The API key you started that server with.
- Admin access to LLM Management in the Barndoor portal.
Step 1: Start your server so Barndoor can reach it
Two defaults bite on nearly every model server, and both must change before Barndoor can use it:- It listens on loopback only, so nothing off that machine can connect.
- It requires no credential, so anything that can connect may use your GPUs freely.
- vLLM
- Ollama
- Anything else
--api-key turns on bearer-token authentication for the /v1 paths. Treat the value as a secret: it is the only thing standing between your GPUs and anyone who can reach the port. Barndoor stores it encrypted and never exposes it to end users.Flags worth knowing about:--served-model-name — control the name clients use
--served-model-name — control the name clients use
vllm serve Qwen/Qwen3-8B reports Qwen/Qwen3-8B in /v1/models and expects that exact string in the model field of a request.--served-model-name qwen3-8b overrides that with something friendlier. Whatever you choose, it must match what you enable in Barndoor in Step 3. The gateway passes the name through untouched.A name without a / is also easier to call directly. The gateway reads everything before the first / in a request’s model field as a provider name, so a model named Qwen/Qwen3-8B can only be called through a Model Route (Step 4) or as <provider name>/Qwen/Qwen3-8B.--enable-auto-tool-choice / --tool-call-parser — tool calling
--enable-auto-tool-choice / --tool-call-parser — tool calling
tool_calls when the server was started for it, and the correct parser depends on the model family:tools through unchanged either way — it can’t compensate for a server that wasn’t started for tool calling.--reasoning-parser — separate reasoning from the answer
--reasoning-parser — separate reasoning from the answer
--reasoning-parser qwen3 (or deepseek_r1, depending on the family) splits thinking tokens into their own field rather than leaving them inline in content. Recommended if your clients render responses directly to users.--max-model-len — cap context length
--max-model-len — cap context length
Making the server reachable
Barndoor connects to your server over the network like any other client, which means the gateway has to be able to route to it. This is independent of which server you run.- Barndoor SaaS
- Self-hosted Barndoor
app.barndoor.ai. You need to give it a routable address:- Put it behind your own load balancer or ingress with a public DNS name and TLS.
- Restrict who can reach that address — an IP allowlist for Barndoor’s egress addresses, or mutual TLS at your edge.
--api-keyis authentication, not network isolation.
Step 2: Add the provider
In the Barndoor portal, go to LLM Management → Providers, click Add Provider, and on the Custom Provider card click Add Custom. Don’t pick one of the named vendor cards. That’s the right choice rather than a compromise. The named cards carry facts about a vendor’s hosted endpoint: its URL, the models it serves, its per-token prices. None of those are knowable for a server you run yourself. The Add Custom Provider form asks for the things that actually vary and assumes nothing else, which is why it works the same for every server in Step 1.How Barndoor checks the connection
On save, Barndoor runs a connectivity check: aGET {base_url}/v1/models with your key, which times out after 10 seconds. The result decides the status shown under the provider’s name in the Providers list:
Step 3: Enable the models
Right after you save the provider, the Enable Models dialog opens. A custom provider has no catalog to pick from, which is the honest state of affairs: only your server knows what it loaded. Type your model name into Custom Models (optional). You can enter several names, separated by commas or new lines. If you skip this step, open the provider later and click Add Models. The name must match what/v1/models reports, exactly. Take it from the server rather than from memory:
Qwen/Qwen3-8B) unless you set --served-model-name, in which case it’s whatever you chose. Copy it character for character — Qwen/Qwen3-8B and qwen3-8b are different models as far as the gateway is concerned, and a mismatch surfaces only when the first request 404s.
Models added this way carry a Custom badge. That’s provenance, not a warning: you typed the name rather than picking it from the Barndoor catalog. It’s also the first place to look when requests to a model fail, since a typo looks exactly like a real model until the first request.
Model '…' has no pricing. Either add a rate under Model Pricing first, or turn off Calculate token cost on the provider. See Cost reporting.Step 4: Create a model route
A model you just enabled is callable only in<provider name>/<model> form, which ties clients to this one provider. Under LLM Management → Model Routes, click Create Route, enter the Route Name your developers will actually use (say qwen3), and point it at the model you just enabled. Developers then call it as qwen3.
This indirection is what lets you move traffic later without touching a single client: repoint the route at a bigger GPU node, or add a hosted provider as a second target so requests spill over when your own capacity is saturated.
Step 5: Verify the connection
Test in layers. Each one isolates a different failure, and a green result at one layer tells you nothing about the next — so resist skipping ahead when something breaks. Almost every failed setup is diagnosed by finding the lowest layer that fails.Your server is serving — on the server host
It's reachable from somewhere else
Barndoor can reach it
The model is exposed to a caller
barndoor.provider field. This is a different question from the previous layer: it proves the route resolves and that model access policy lets this caller use it.If the provider is healthy but the model isn’t listed here, the problem is the route or the access policy, not your server. Go back to Step 4, and check LLM Management → Model Access.A request completes — both ways
"stream": true added. Streaming goes through a different path, an SSE relay with its own idle timeout, and it’s what coding assistants and chat UIs actually use. A passing non-streaming request is not evidence that streaming works.data: chunks, a final chunk carrying usage, then data: [DONE]. The usage chunk is what makes a streamed request countable. The gateway asks for it on every streamed request even when the client doesn’t, so your server must accept stream_options. If the usage chunk is missing, tokens won’t reach your reports.Governance actually applied
Cost reporting for self-hosted models
Self-hosted inference has no market rate (the cost is your own GPU time), so Barndoor ships no default pricing for self-hosted models. Until you decide, usage reporting counts tokens accurately, reports the cost as zero, and shows the model as Unpriced. You have two ways to make that a decision rather than a gap:- Record self-hosted traffic as not metered. Edit the provider, turn off Calculate token cost, and set Billing arrangement to Self-hosted. Tokens are still counted, cost is recorded as $0 on purpose, and the model shows Not metered instead of Unpriced. The Billing arrangement is required when cost calculation is off.
- Give it a rate. Leave Calculate token cost on and set your own rates under LLM Management → Model Pricing. A reasonable approach is to divide the fully loaded hourly cost of the instance by the tokens it produces in that hour, and enter the result as an input and output rate. It doesn’t need to be exact to be useful: even a rough number makes “what did this team actually consume” answerable, and lets a cost budget act as a real ceiling. See Managing Model Pricing.
Troubleshooting
could not reach upstream provider
could not reach upstream provider
- vLLM is bound to loopback — restart it with
--host 0.0.0.0. - A firewall or security group blocks the port from Barndoor’s side.
- You’re on Barndoor SaaS and the address is private. See Making the server reachable.
- You’re using
https://with a self-signed or private-CA certificate.
connection to upstream timed out after 10s
connection to upstream timed out after 10s
upstream rejected the credentials (HTTP 401)
upstream rejected the credentials (HTTP 401)
base_url must not end in '/v1'
base_url must not end in '/v1'
/v1. Use the corrected URL the message suggests, which is the same address without the trailing /v1. The gateway adds /v1 itself.Provider shows Enabled, but requests fail with 404
Provider shows Enabled, but requests fail with 404
/v1 paths, and whether the base URL points at the host root rather than a sub-path.Requests fail with an unknown-model error
Requests fail with an unknown-model error
curl http://<host>:8000/v1/models: the id there is the name to enable, letter for letter. If you added --served-model-name after configuring Barndoor, the enabled name is now stale.If the gateway itself answers 404 model '…' not found, the problem is the name the client sent, not the server. A model ID containing / (such as Qwen/Qwen3-8B) can’t be called by its bare name. Call it through a Model Route, or as <provider name>/Qwen/Qwen3-8B.Tool calls come back as prose
Tool calls come back as prose
--enable-auto-tool-choice and a --tool-call-parser matching the model family. Barndoor forwards the tools array unchanged; it cannot synthesize structured tool calls from a server that isn’t producing them.Current limits
- Chat completions and streaming are the supported surface. Claude Code and other Anthropic-format clients can use a self-hosted model too, because the gateway translates Messages requests into chat completions. The gateway also forwards embeddings, legacy completions, and Responses API requests to your server unchanged, but those aren’t verified on this path and work only if your server implements them.
- Tool calling and reasoning output depend entirely on how you started the server, so Barndoor makes no promise about them on your behalf. They pass through when your server produces them.
- Barndoor SaaS cannot reach a server on a private network. There is no tunnel or agent for this today; the server needs a routable address, or Barndoor needs to run inside your infrastructure.