Skip to main content

Overview

Once your LLM Gateway is set up — see Using the LLM Gateway — the Controls group of LLM Management is where admins govern how the people, keys, and agents in their organization consume it: capping spend, throttling busy callers, and deciding who may use which models.
The Controls group of the LLM Management hub showing the Budgets, Rate Limits, and Model Access sections
The Controls group has three sections:
Per-request usage and cost reporting lives separately, under Reporting → LLM Usage Dashboard — see Reading the LLM Usage Dashboard.

Before You Begin

You’ll need:
  • A Barndoor account with admin privileges.
  • An LLM Gateway with at least one provider and one model route — see Using the LLM Gateway.
  • Nothing extra for spending budgets: Barndoor prices requests from its managed default catalog automatically. Add your own prices under LLM Management → Model Pricing only to override those defaults or to price a model the catalog doesn’t cover. See Managing Model Pricing.

Key Concepts

Budgets and rate limits share the same two independent dimensions. Model access policies use the first one.

Scope: who a rule applies to

For every scope except Organization you can pick one specific entity, or leave the picker on All users / All groups / etc. Leaving it blank does not create one shared pot:
  • User, API Key, AI Agent left blank: each user, key, or agent gets their own allowance equal to the limits you set.
  • Group, Role left blank: each group or role gets one allowance, shared by everyone in it.
A User budget or rate limit left on All users can additionally be narrowed with Limit to group. Every member of that group then gets their own allowance, and anyone who joins the group picks it up on their next request with no admin action. Compare this with a Group scope, which pools one allowance across all members.

Target: what a rule counts

By default a budget or rate limit counts All targets. Set a Target to count only matching traffic: Scope and Target combine: a rule scoped to User: Ada with Target LLM Provider: OpenAI counts only Ada’s traffic that OpenAI serves.
A Model Route or Routing Policy target matches the name the caller sent in model — before the gateway routes it — not the model it ended up on. A routing policy target therefore counts every request for that policy, pooled across all of its slots. To cap a specific underlying model regardless of which route reached it, target LLM Provider with an Upstream model instead.
Routing policies appear in the target picker only when your organization has routing policies enabled. That feature may not be enabled for your organization yet — contact Barndoor. See Routing Policies.

How Limits Compose

Governance runs on /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/responses, and /v1/messages. (/v1/models and /v1/messages/count_tokens are not metered.) Checks run in a fixed order; the first denial stops the request.
Inbound request
Every request
1
API-key authThe bd-… key is valid, active, and not revoked
401
2
Rate limits with no TargetRequests and tokens per window
429
3
Budgets with no TargetToken budgets, then spending budgets
429
4
Model resolutionThe model maps to a route or provider model
404
5
Model access policiesAllowlists and denylists for this caller
403 · 503
For each route in the plan
6
Rate limits with a TargetMatching routes are dropped
Drop route
7
Blocking budgets with a TargetMatching routes are dropped
Drop route
Forward
Send to the first remaining routeUsage is recorded from the tokens the provider reports
429 if none left
4xx Request refusedDrop route Route skipped, others still tried
Things worth knowing about this chain:
  • Usage is recorded after the upstream responds, from the token counts the provider reports. Tokens counted are input (including cache reads and cache writes) plus output (which includes reasoning tokens). The pre-check enforces the limit; the recording keeps the counter accurate.
  • Every applicable rule is checked, and the most restrictive result wins. A user covered by an org budget and a user budget is checked against both, and both are debited.
  • A more specific rule overrides a broader one of the same scope type and target. For example, a User: Ada budget replaces the All users budget for Ada alone, and a Limit to group budget replaces the All users budget for that group’s members. This lets you raise one person’s or one group’s allowance above the default. Rules of different scope types (say, Organization and User) don’t override each other — both apply.
  • Rules with a Target fail over instead of hard-blocking. When a targeted budget or rate limit is exhausted, the gateway removes just the matching routes from the request’s route plan and tries the remaining ones. The caller sees a 429 only when every eligible route is blocked. See Failover, Cooldowns, and Route Health.
  • Controls fail open on a transient outage. If the usage store can’t be reached within 2 seconds, budgets and rate limits are skipped for that request rather than blocking live traffic. Model access is the exception — see Model Access.

Budgets

Budgets cap how much a scope may consume in a calendar period. A budget can cap tokens, spending, or both — whichever limit is hit first triggers the action.
LLM Budgets section listing budgets with token and spending usage bars
The LLM Budgets list shows each budget’s Usage (a TOKENS and/or COST bar), Status, Period, Resets in, On exhaust, Created, Last Updated, and an Enabled toggle. Usage refreshes about every 10 seconds. You can search by name, scope, target, or period, and sort by any header; the default sort is Usage, highest first, so budgets closest to exhaustion float to the top.
  • Status shows an amber NN% pill once usage crosses one of the budget’s alert thresholds, and a red Exhausted pill at 100%.
  • Per-member budgets (a blank entity at any scope other than Organization) have no single usage figure. Their row reads Per member — expand for usage; click the chevron at the left of the row to see each member’s usage and breach status.

Creating a budget

1

Open LLM Management → Budgets and click Create Budget

2

Name, period, and scope

  • Name — shown in the denial message callers receive and in alerts, so make it recognizable (for example Engineering monthly).
  • Period — Daily, Weekly, or Monthly (default).
  • Scope — see Scope. Pick an entity, or leave it on All … for per-member allowances. For a User scope left on All users, optionally set Limit to group.
3

Action on Exhaust and alerts

  • Action on Exhaust — Block Requests (default) returns 429 once the limit is reached. Warn Only keeps serving requests and only raises alerts.
  • Alert Thresholds (%) — comma-separated percentages of the limit at which to raise an alert, for example 80, 90 (the default). Leave blank for none.
4

Target (optional)

Leave on All targets, or narrow the budget as described in Target. Only Block budgets remove targeted routes; a Warn Only budget with a Target only alerts.
5

Budget type and limits

  • Budget Type — Tokens Only, Spending ($) Only (default), or Both (first hit).
  • Token Limit — total tokens allowed in the period.
  • Spending Limit ($) — a USD cap, costed from model pricing.
  • Reason for change (optional) — a note recorded in the audit trail.
6

Click Create

Create Budget dialog with name, period, scope, action, target, and limits filled in
When editing a budget you can change its name, period, action, budget type, limits, and thresholds. Scope and Target can’t be changed after creation — delete the budget and create a new one.
A spending limit on a budget targeted at a provider whose Calculate token cost setting is off never fills up, because that provider’s requests are recorded at $0. The dialog warns you when this applies; use a token limit instead.

How budgets are counted and reset

  • Tokens — input tokens (including cache reads and writes) plus output tokens, as reported by the provider.
  • Spending — the request’s cost from Model Pricing, including cache and long-context rates. Requests that cost $0 (unpriced models, providers that don’t calculate token cost) don’t move a spending counter.
  • Reset — every period boundary is UTC, with no per-organization override: Daily at 00:00 UTC, Weekly on Monday at 00:00 UTC, Monthly on the 1st at 00:00 UTC. The Resets in column shows the countdown, and its tooltip gives the exact UTC instant.

Alerts

Barndoor raises admin alerts from each budget’s own thresholds:
  • Threshold alert (warning) — when usage crosses one of the configured percentages.
  • Exhausted alert (critical) — on a Block Requests budget, as soon as the request that crosses 100% completes. That request is still allowed to finish; the next one is blocked.
Token and spending limits alert separately. Each alert is sent at most once per day per budget and threshold, and on a per-member budget, once per member, naming the member. Alerts link straight to the budget in the list (and open the member’s row on a per-member budget).

When a caller hits a budget

A Block Requests budget with no Target returns:
A spending budget reads Spending monthly budget exhausted (budget: Engineering monthly) (100% used). When a budget with a Target has removed every eligible route, the 429 has the same type and a message naming each blocked route, for example:
tokens_used is often slightly above the limit in the message. The gateway checks the counter before forwarding and debits actual usage after the provider responds, so the request that crosses the limit finishes and the next one is denied. The percentage is rounded, so 1,001,234 / 1,000,000 prints as 100%.
Tokens give predictability; spending tracks dollars across models with different prices. Both (first hit) is the safest configuration: the token cap absorbs a price change, and the spending cap catches an expensive model nobody noticed.

Rate Limits

Rate limits cap throughput on a rolling 60-second window. Use them to smooth bursts and stop one noisy caller or runaway loop from crowding out everyone else. (Use budgets to control total consumption over a day, week, or month.)
LLM Rate Limits section listing policies with live RPM and TPM usage
The LLM Rate Limits list shows Live (60s) usage, Scope, Req/min, Tokens/min, Traffic, Last Updated, and Enabled, refreshed about every 10 seconds. The default sort is Live (60s), so policies closest to their limit surface first. A policy that gives each caller their own window (a blank entity) has a chevron to expand per-caller usage.

Creating a rate limit

1

Open LLM Management → Rate Limits and click Create Policy

2

Name and scope

  • Name — appears in the denial message.
  • Scope — see Scope, including Limit to group for a User scope left on All users.
3

Target (optional)

Leave on All targets, or narrow the policy as described in Target. While a targeted limit is exceeded, only the matching routes stop serving.
4

Limits

Set Requests / min, Tokens / min, or both. At least one is required.
5

Click Create

Create Rate Limit Policy dialog with name, scope, target, and limits
Unlike budgets, a rate limit’s scope and target can be edited later. (The one exception: a policy using Limit to group has its scope locked.)
Older policies may show a scope of Model or LLM Provider. Those scopes never enforced correctly, and the edit dialog says so. Recreate them with the identity as Scope and the model route or provider as Target.

When a caller hits the limit

A targeted rate limit that blocks every route returns 429 rate_limit_error with Retry-After: 60 and a message like All routes for model 'gpt-5.5' are blocked by rate limits: ….
Retry-After is always the full window length, 60 seconds: waiting that long is guaranteed to clear the window, though callers often succeed sooner as older requests age out. The OpenAI and Anthropic SDKs honor it automatically.
Requests per minute and tokens per minute are enforced independently. A caller under the request cap can still be throttled by the token cap, and vice versa. Token usage is recorded after each response, so a single very large request can overshoot the token limit before the next one is refused.

Model Access

Model access policies decide whether a caller may use a route at all. The section has two parts: an org-wide posture and a list of Model Access Policies.
Model Access section with the posture card above the list of allowlist and denylist policies

Posture: open or locked down

The card at the top of the section, Require an explicit grant to use a model, sets what happens to a route that no policy mentions: Turning it on asks for confirmation, and warns you if no allowlist exists yet — in that state every request would be denied, including your own.
Allowlists only ever narrow access in the Open posture — they never grant anything beyond it. To make “no grant, no access” the rule, switch the posture on.

Creating a policy

1

Open LLM Management → Model Access and click Add Policy

2

Name, policy type, and scope

  • Name — appears in the denial message.
  • Policy Type — Allowlist (only these targets are allowed) or Denylist (these targets are blocked).
  • Scope — Organization, Group, Role, User, API Key, or AI Agent. Leaving the entity blank applies the policy to every entity of that type.
3

Add targets

Pick a kind under Add a target, choose a value, and click Add target. Mix any number of targets:Model, route, and provider + model values accept a trailing * wildcard (for example claude-*).
4

Save

The policy appears in the list. Use its Enabled toggle to switch it off without deleting it.
Create Model Access Policy dialog with a Route Group target added

How policies combine

Model access has no “more specific overrides broader” rule — every policy whose scope matches the caller is evaluated together:
  1. If any matching denylist matches the route, the route is denied.
  2. Otherwise, if any allowlist applies to the caller, at least one of them must grant the route.
  3. Otherwise the posture decides: allowed when Open, denied when Locked down.
Policies are evaluated per route. If a route name fans out to several providers and only some are denied, the request proceeds on the allowed ones — a denied provider is never used, not even for failover. The caller sees a 403 only when no route survives.
Provider targets are deliberately asymmetric. In a denylist, a provider covers everything routed to it, including custom route names. In an allowlist, it grants only the provider’s own models — to grant a custom route name that points at the provider, add it as a Model Route or Provider + Model target.

When a caller is denied

The message depends on why the route was refused:
Model access fails closed. If Barndoor can’t read the organization’s posture when a decision depends on it, the request is refused with 503 service_unavailable_error (“Model access posture for this organization is temporarily unavailable…”), never silently allowed. A matching denylist or granting allowlist is still honored during such an outage.

Governance settings

Three org-wide switches sit alongside the controls. Each takes effect on the next request.

What callers see when something is denied

Denials use the OpenAI-style envelope { "error": { "message", "type", "code" } }, so common LLM SDKs surface them as ordinary errors. On /v1/messages, rate-limit denials use Anthropic’s envelope instead: { "type": "error", "error": { "type", "message" } }.

Troubleshooting

Read error.message: it names the rule — (policy: …) for a rate limit, (budget: …) for a budget, or a list of blocked routes for a targeted rule. Open that rule to check its configuration and live usage. For a per-member rule, expand the row to find the user.
Budget periods roll over at midnight UTC, not local time: daily at 00:00 UTC, weekly on Monday, monthly on the 1st. The Resets in column’s tooltip shows the exact UTC instant.
A Group scope pools one allowance across all members. To give each member their own, use Scope: User, leave it on All users, and set Limit to group to that group.
Agent-scoped rules only match requests made with an API key bound to that agent. Traffic on an unbound key isn’t counted.
Check the policy’s scope and that it’s enabled. Remember that in the Open posture an allowlist only restricts the callers it applies to — everyone else is unaffected. Wildcards must be trailing: claude-* works, *-mini doesn’t.
Either the models being called have no price (neither your rules nor Barndoor’s defaults cover them), or the budget targets a provider whose Calculate token cost setting is off. See Managing Model Pricing.
Set Alert Thresholds (default 80, 90) to be alerted before the budget blocks anyone. The Status column shows breached thresholds at a glance, and Reporting → LLM Usage Dashboard shows consumption trends.

Frequently Asked Questions

Use a budget for total consumption over a long horizon (a team’s monthly spend). Use a rate limit for short-term load (smoothing bursts, stopping a runaway loop). They compose: a budget for how much, a rate limit for how fast.
Yes. Create the default as a User rule on All users, then a second rule for User: that person with the same Target and a higher limit. The specific rule replaces the default for that user only. The same works for a group of users via Limit to group.
Yes. Embedding requests count tokens against budgets and rate limits and are subject to model access, like chat requests.
Yes — they record usage and raise threshold alerts without blocking anyone. A common pattern is to run a budget in Warn Only for a period or two to calibrate the limit, then switch it to Block.
Scope it tightly — a single test User or a dedicated API Key — and call the gateway with that user or key. Widen the scope once you’re happy with the behavior.

Need Help?

Reach out to [email protected] with:
  • The name of the budget, rate limit, or policy involved.
  • The exact error.message returned to the caller (if any).
  • The scope and target of the rule.