smart — in the model field, and for each request the gateway picks one of the policy’s slots: an ordered list of models from least to most capable. Simple requests go to a cheap slot; harder ones go higher.
The choice is made by a determiner, a small, fast model you pick that reads the request and returns a slot. Routing rules let you constrain that choice in plain English — for example, “contract review never goes below Sonnet”.
Routing policies replace the earlier Smart Models feature. Existing smart models carried over as routing policies with the same name, so callers don’t need to change anything.Routing Policies and Routing Rules may not be enabled for your organization yet — contact Barndoor if you don’t see them under LLM Management.
Key terms
Create a routing policy
1
Open LLM Management → Routing Policies
Click Add Routing Policy.
2
Name it
Routing policy name is the value callers put in
model (for example smart). It can’t be the same as an existing model route’s name. Description is optional. Leave Enabled on.3
Pick a posture
Savings prefers the cheaper slot when the choice is close, Quality prefers the stronger one, and Balanced (the default) has no preference. Posture only breaks ties — it never overrides a model access limit or a rule.
4
Fill the model slots
A policy needs at least two slots, ordered least → most capable (the form starts with Fast, Standard, and Premium). For each slot:
- Label — a name for the slot.
- Target — switch between Route (one of your model routes, which keeps that route’s failover) and Model (a specific provider model). The target must have at least one enabled route, and can’t be the policy itself.
- Routing guidance (optional) — what this slot is best for, for example “code” or “long-form writing”. The determiner sees this next to the slot.
5
Configure the determiner
- Determiner model — the model or route that picks a slot. It must support JSON-mode output (
response_format: json_object). A small, fast model is usually right. - Characters of the request it reads — how much of the request is sent to the determiner (default 12,000). Lower is cheaper and faster; too low and it judges on an unrepresentative excerpt.
- Determiner instructions (optional) — replaces the default instructions. The slot list, request details, and required JSON output format are always added automatically.
- Fallback slot — the slot used when the determiner errors or returns no usable answer (default: slot 1).
6
Test it, then save
Under Test routing, enter a sample request and click Preview routing. The result shows the slot and model it would get, what decided it, and the determiner’s reason — without saving anything. If you change the configuration afterwards, the result is marked stale. Click Save.
How a slot is chosen
For each request to a routing policy, the gateway:- Estimates the request’s size from all its text, and sets a minimum slot from your context thresholds.
- Skips the determiner when the answer is already clear. A request containing images goes to the top slot. A request at or above the last context threshold also goes to the top slot.
- Asks the determiner. It sees the latest user message only — not the whole conversation, system prompt, or tool definitions — cut to the character limit you set. (Claude Code’s injected
<system-reminder>context is removed first, so it doesn’t make every request look complex.) Along with it, the determiner sees the slots it may choose from, their guidance, the policy’s enabled rules, whether the request uses tools or images, and the posture. - Never goes below the minimum slot, whatever the determiner picked. If the determiner fails, the Fallback slot is used (still no lower than the minimum).
- Applies model access and rules — see Constraints.
- Sends the request to the chosen slot’s target, which fails over across its own targets as usual.
Constraints: model access and rules
The determiner’s pick can still be moved after it is made. None of these ever grant access — they only narrow what a request reaches.- Model access. Slots the caller isn’t allowed to use are left out of what the determiner sees. If it still lands on one, the request moves up to the next slot the caller may use. If the caller may use none of the policy’s slots, the request is refused with a
403. - Routing rules. If the determiner says the request matches a rule, the gateway looks that rule up itself and applies its floor and bans. A ban removes a slot; a floor raises the pick to at least that slot. When the two can’t both be met, the ban wins. If a rule’s floor is above anything the caller may use, they get the best slot they can reach rather than an error. If a ban removes every slot the caller could use, the request is refused with a
403. - Size and images. The minimum slot from step 1 is never traded away. If the caller isn’t allowed any slot at or above it, the request is refused with a
403asking an administrator to widen their model access, rather than answered by a model too small for it.
Conversations stay on one model
Switching models mid-conversation makes agents incoherent and discards the provider’s prompt cache, so the gateway pins a conversation to the slot it was first routed to. A conversation is identified by the organization, the API key, and the text of its first user message. A pin lasts 55 minutes and is refreshed on every turn that uses it. On each later turn:Two conversations under the same API key that open with exactly the same first message share a pin. This mostly affects automated agents that always start with the same boilerplate.
Routing rules
Rules are written per policy under LLM Management → Routing Rules. Choose the policy from Routing policy, then click Add Rule.
A rule needs a floor, a ban, or both — one with neither would never change anything. A floor or ban is stored as a slot position, so if you later remove slots from the policy, a floor or ban on a slot that no longer exists stops applying (it’s dropped, not moved to the new top slot).
A rule conflicts when its floor can never be met — because it bans every slot at or above its own floor, or because another rule bans them. Saving still succeeds, but the tab shows a warning banner and the rule carries a Conflict badge. If both rules match a request, the ban wins. Deleting a rule tells you which slots its matching requests may be routed to again.
Test a request
The Test a request panel runs sample text through the selected policy’s determiner and rules. The result shows the slot and model the request would get, Rule: name or No rule matched, and — when something moved the determiner’s pick — what it picked and why it moved (a model access policy, a tools or images requirement, the rule’s floor, or the rule’s ban). Sample text is never stored or logged.Require requests to use a routing policy
By default, callers can still name a model or model route directly, which bypasses routing policies entirely. To stop that, turn on Require requests to use a routing policy at the top of Routing Policies and confirm Require routing policies. The badge changes from Optional to Required. While it’s on:- A request must name an enabled routing policy. Naming a model (
anthropic/claude-…) or a model route is refused with a400that lists the policies available — never silently routed somewhere else. - Integrations that name models directly will start failing, so turn this on only after they’ve moved to a policy name.
Budgets, model access, and cost
- Model access is checked against both the policy name the caller sent and the model the policy chose. To allow or block a policy itself, target it with Model Route or Routing Policy in Model Access.
- Budgets and rate limits targeted at Model Route or Routing Policy match the name the caller sent — the policy’s name — not the slot it was routed to. A budget on
smartcounts everything sent tosmart, whichever model served it. To cap one underlying model however it’s reached, target its Upstream Model instead. See LLM Controls. - The determiner call is billed to the caller. It counts against their budgets and rate limits and is recorded under the determiner model’s name, not the policy’s. On the LLM Usage Dashboard, grouping by Routing Role separates User requests from Routing overhead (determiner calls) and Routing preview (admin) (determiner calls made by Preview routing and Test a request). Preview spend is recorded but never counts against anyone’s budget or rate limit.
Seeing why a request went where it did
Response headers. Responses to a routing-policy request includex-bd-router-* headers, among them:
Audit History. Routed requests can be filtered by:
Filtering Routing Override to anything other than Not overridden finds every request where a limit or rule changed the model.