> ## Documentation Index
> Fetch the complete documentation index at: https://docs.barndoor.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Routing Policies

> Give callers one model name — like smart — and let the gateway pick the cheapest model strong enough for each request, constrained by plain-English routing rules, model access, and budgets.

A **routing policy** is a model name that doesn't point at one model. Callers put its name — say `smart` — in the `model` field, and for each request the gateway picks one of the policy's **slots**: an ordered list of models from least to most capable. Simple requests go to a cheap slot; harder ones go higher.

The choice is made by a **determiner**, a small, fast model you pick that reads the request and returns a slot. **Routing rules** let you constrain that choice in plain English — for example, "contract review never goes below Sonnet".

<Note>
  Routing policies replace the earlier **Smart Models** feature. Existing smart models carried over as routing policies with the same name, so callers don't need to change anything.

  **Routing Policies** and **Routing Rules** may not be enabled for your organization yet — contact Barndoor if you don't see them under **LLM Management**.
</Note>

## Key terms

| Term | Meaning |
| - | - |
| **Slot** | One position in the policy's ordered list. Slot 0 is the least capable (and usually cheapest); the last slot is the most capable. Each slot targets a model route or a specific provider model. |
| **Determiner** | The model that reads each request and picks a slot. Its call is a real, billed LLM request. |
| **Context threshold** | A request size, in estimated tokens, at which requests start at a higher slot regardless of what the determiner says. |
| **Posture** | Which way the determiner leans when two slots are about equally suitable. |
| **Routing rule** | A plain-English description of a kind of request, plus a floor ("never below") and/or bans ("never use"). |

## Create a routing policy

<Steps>
  <Step title="Open LLM Management → Routing Policies">
    Click **Add Routing Policy**.
  </Step>

  <Step title="Name it">
    **Routing policy name** is the value callers put in `model` (for example `smart`). It can't be the same as an existing model route's name. **Description** is optional. Leave **Enabled** on.
  </Step>

  <Step title="Pick a posture">
    **Savings** prefers the cheaper slot when the choice is close, **Quality** prefers the stronger one, and **Balanced** (the default) has no preference. Posture only breaks ties — it never overrides a model access limit or a rule.
  </Step>

  <Step title="Fill the model slots">
    A policy needs at least two slots, ordered least → most capable (the form starts with **Fast**, **Standard**, and **Premium**). For each slot:

    * **Label** — a name for the slot.
    * **Target** — switch between **Route** (one of your [model routes](/how-tos/llm-gateway-failover-and-cooldowns), which keeps that route's failover) and **Model** (a specific provider model). The target must have at least one enabled route, and can't be the policy itself.
    * **Routing guidance** (optional) — what this slot is best for, for example "code" or "long-form writing". The determiner sees this next to the slot.

    Between each pair of slots, **Requests of at least N tokens start at** *next slot* sets a context threshold. Thresholds must increase down the list (defaults: 8,000 and 32,000).
  </Step>

  <Step title="Configure the determiner">
    * **Determiner model** — the model or route that picks a slot. It must support JSON-mode output (`response_format: json_object`). A small, fast model is usually right.
    * **Characters of the request it reads** — how much of the request is sent to the determiner (default 12,000). Lower is cheaper and faster; too low and it judges on an unrepresentative excerpt.
    * **Determiner instructions** (optional) — replaces the default instructions. The slot list, request details, and required JSON output format are always added automatically.
    * **Fallback slot** — the slot used when the determiner errors or returns no usable answer (default: slot 1).
  </Step>

  <Step title="Test it, then save">
    Under **Test routing**, enter a sample request and click **Preview routing**. The result shows the slot and model it would get, what decided it, and the determiner's reason — without saving anything. If you change the configuration afterwards, the result is marked stale. Click **Save**.
  </Step>
</Steps>

Each saved policy appears as a card showing its posture, whether it's enabled, the determiner, the slots, and the context thresholds.

<Warning>
  Deleting a policy asks you to type its name to confirm. Clients calling that name stop resolving, and the policy's routing rules are deleted with it. Disabling a policy is not a soft delete either: requests to a disabled policy are **refused** with a `400`, rather than served by some other model with the same name.
</Warning>

## How a slot is chosen

For each request to a routing policy, the gateway:

1. **Estimates the request's size** from all its text, and sets a minimum slot from your context thresholds.
2. **Skips the determiner when the answer is already clear.** A request containing images goes to the top slot. A request at or above the last context threshold also goes to the top slot.
3. **Asks the determiner.** It sees the latest user message only — not the whole conversation, system prompt, or tool definitions — cut to the character limit you set. (Claude Code's injected `<system-reminder>` context is removed first, so it doesn't make every request look complex.) Along with it, the determiner sees the slots it may choose from, their guidance, the policy's enabled rules, whether the request uses tools or images, and the posture.
4. **Never goes below the minimum slot**, whatever the determiner picked. If the determiner fails, the **Fallback slot** is used (still no lower than the minimum).
5. **Applies model access and rules** — see [Constraints](#constraints-model-access-and-rules).
6. **Sends the request to the chosen slot's target**, which fails over across its own targets as usual.

### Constraints: model access and rules

The determiner's pick can still be moved after it is made. None of these ever grant access — they only narrow what a request reaches.

* **Model access.** Slots the caller isn't allowed to use are left out of what the determiner sees. If it still lands on one, the request moves **up** to the next slot the caller may use. If the caller may use none of the policy's slots, the request is refused with a `403`.
* **Routing rules.** If the determiner says the request matches a rule, the gateway looks that rule up itself and applies its floor and bans. A ban removes a slot; a floor raises the pick to at least that slot. When the two can't both be met, **the ban wins**. If a rule's floor is above anything the caller may use, they get the best slot they can reach rather than an error. If a ban removes every slot the caller could use, the request is refused with a `403`.
* **Size and images.** The minimum slot from step 1 is never traded away. If the caller isn't allowed any slot at or above it, the request is refused with a `403` asking an administrator to widen their model access, rather than answered by a model too small for it.

The determiner can only *name* a rule; the floor and bans always come from the rule as you saved it. A determiner that names a rule that doesn't exist, or a disabled one, applies nothing.

### Conversations stay on one model

Switching models mid-conversation makes agents incoherent and discards the provider's prompt cache, so the gateway **pins** a conversation to the slot it was first routed to. A conversation is identified by the organization, the API key, and the text of its first user message. A pin lasts 55 minutes and is refreshed on every turn that uses it.

On each later turn:

| Turn | What happens |
| - | - |
| A tool result coming back | Stays on the pinned model |
| A short acknowledgement ("yes", "go ahead", "thanks") | Stays on the pinned model |
| Newly needs tools or images, or grows past a context threshold | Re-routed, and only ever to a **stronger** model |
| Already on the strongest model the caller may use, and no rule on the policy bans anything | Stays on the pinned model |
| Anything else | The determiner runs again, so a rule can apply to something pasted mid-conversation (for example a contract in the sixth message). This re-check can move the conversation to a different model, including a cheaper one — for example when a rule bans the model it was on. |

<Note>
  Two conversations under the same API key that open with exactly the same first message share a pin. This mostly affects automated agents that always start with the same boilerplate.
</Note>

## Routing rules

Rules are written per policy under **LLM Management → Routing Rules**. Choose the policy from **Routing policy**, then click **Add Rule**.

| Field | What it does |
| - | - |
| **Name** | How the rule is shown (up to 120 characters). |
| **When does this apply?** | A plain-English description of the requests it covers (up to 2,000 characters). This is what the determiner matches against — describe the requests, not the rule. No example prompts needed. |
| **Never below** (optional) | The least capable slot ever acceptable for matching requests. |
| **Never use** (optional) | Slots matching requests must never go to. |
| **Enabled** | A disabled rule is ignored. |

A rule needs a floor, a ban, or both — one with neither would never change anything. A floor or ban is stored as a slot position, so if you later remove slots from the policy, a floor or ban on a slot that no longer exists stops applying (it's dropped, not moved to the new top slot).

A rule conflicts when its floor can never be met — because it bans every slot at or above its own floor, or because another rule bans them. Saving still succeeds, but the tab shows a warning banner and the rule carries a **Conflict** badge. If both rules match a request, the ban wins. Deleting a rule tells you which slots its matching requests may be routed to again.

### Test a request

The **Test a request** panel runs sample text through the selected policy's determiner and rules. The result shows the slot and model the request would get, **Rule: *name*** or **No rule matched**, and — when something moved the determiner's pick — what it picked and why it moved (a model access policy, a tools or images requirement, the rule's floor, or the rule's ban). Sample text is never stored or logged.

<Tip>
  **Preview routing** in the policy editor tests an unsaved configuration, so it doesn't apply the policy's rules. To see rules take effect, use **Test a request** on the Routing Rules tab.
</Tip>

## Require requests to use a routing policy

By default, callers can still name a model or model route directly, which bypasses routing policies entirely. To stop that, turn on **Require requests to use a routing policy** at the top of **Routing Policies** and confirm **Require routing policies**. The badge changes from **Optional** to **Required**.

While it's on:

* A request must name an **enabled** routing policy. Naming a model (`anthropic/claude-…`) or a model route is refused with a `400` that lists the policies available — never silently routed somewhere else.
* Integrations that name models directly will start failing, so turn this on only after they've moved to a policy name.

<Warning>
  If there are no enabled routing policies, every request would be refused. The confirmation warns you, and deleting the last enabled policy while this is on warns you again.
</Warning>

## Budgets, model access, and cost

* **Model access** is checked against both the policy name the caller sent *and* the model the policy chose. To allow or block a policy itself, target it with **Model Route or Routing Policy** in [Model Access](/how-tos/use-llm-controls).
* **Budgets and rate limits** targeted at **Model Route or Routing Policy** match the name the caller sent — the policy's name — not the slot it was routed to. A budget on `smart` counts everything sent to `smart`, whichever model served it. To cap one underlying model however it's reached, target its **Upstream Model** instead. See [LLM Controls](/how-tos/use-llm-controls).
* **The determiner call is billed to the caller.** It counts against their budgets and rate limits and is recorded under the determiner model's name, not the policy's. On the [LLM Usage Dashboard](/how-tos/read-llm-usage-dashboard), grouping by **Routing Role** separates **User requests** from **Routing overhead** (determiner calls) and **Routing preview (admin)** (determiner calls made by **Preview routing** and **Test a request**). Preview spend is recorded but never counts against anyone's budget or rate limit.

## Seeing why a request went where it did

**Response headers.** Responses to a routing-policy request include `x-bd-router-*` headers, among them:

| Header | Meaning |
| - | - |
| `x-bd-router-model` | The policy name the caller sent |
| `x-bd-router-target-model` | The model or route it was sent to |
| `x-bd-router-slot-index` | The slot served (0 = least capable) |
| `x-bd-router-source` | What decided: `heuristic`, `determiner`, `default_on_failure`, `sticky_stay`, or `pin_break` |
| `x-bd-router-reason` | A short explanation, written by the determiner or the gateway |
| `x-bd-router-clamp-reason` | What moved the pick, if anything: `none`, `model_access`, `capability_floor`, `rule_floor`, or `rule_ban` |
| `x-bd-router-picked-before-clamp` | The slot the determiner picked, when something moved it |
| `x-bd-router-matched-rule-name` | The rule that constrained the request, if any |
| `x-bd-router-skipped-reason` | On a pinned turn, why the determiner didn't run: `tool_turn`, `at_max_slot`, or `no_new_content` |

<Warning>
  The determiner writes `x-bd-router-reason` about the request it read, so it can quote the request's text.
</Warning>

**Audit History.** Routed requests can be filtered by:

| Filter | Values |
| - | - |
| **Routing Decision** | **Heuristic** (decided by size or images), **Determiner**, **Determiner failed** (fallback slot used), **Sticky pin** (stayed on the conversation's model), **Pin break** (the conversation was re-routed) |
| **Routing Override** | **Not overridden**, **Access limit**, **Capability required**, **Rule floor**, **Rule ban** |
| **Routing Re-check** | Why a pinned turn skipped the determiner: **Tool turn**, **Highest model reached**, **No new content** |

Filtering **Routing Override** to anything other than **Not overridden** finds every request where a limit or rule changed the model.
