Choosing the window and scope
Everything on the page — the summary cards, every chart and every table — follows the same time window and the same scope filters.Time window
The date button in the header shows the active window and the timezone it’s expressed in. Open it to pick a preset (Last 1 hour through Last 30 days, or Today, Yesterday, Week to date, Month to date, Last month) or drag a custom range on the calendar. The arrows either side step the window back or forward by its own width.- Rolling presets end at “now” and move forward each time the page refreshes. Stepping the window, or picking dates on the calendar, pins it to fixed start and end times instead.
- Picking days on the calendar selects whole days, in the active timezone. Future days can’t be selected.
- The window is part of the URL, so copying the address shares the exact view. A rolling preset stays rolling for whoever opens the link.
Timezone
The dashboard uses UTC by default. Switch between UTC and Local (your browser’s timezone) at the bottom of the date picker. The choice decides where each day, week and month starts, so it changes which period a request lands in — and therefore the per-period figures — but never the window’s totals for the same start and end times.Keep UTC when comparing against provider invoices, which are almost always denominated in UTC days and months. Reset in the date picker restores the default window, granularity and timezone.
Granularity
The Hourly / Daily / Weekly selector sets the length of each period on the charts and in the By period table. It follows the window unless you change it: hourly up to two days, daily up to 90 days, weekly beyond that. Options that would be unreadable for the window are greyed out — hourly over more than eight days, daily under six hours, weekly under 14 days. Weeks start on Monday.Scope filters
Two independent filters narrow the whole dashboard. Each is optional, and when both are set they combine — By Group “Platform” together with By Provider “anthropic” shows only that group’s Anthropic traffic.
Picking a dimension without a value leaves that filter off, so the dashboard never goes blank mid-selection.
By AI Agent matches the agent recorded on each request at the time it was made — not the agent’s currently bound keys — so an agent’s history doesn’t change when keys are unbound or rebound. It only appears once API keys can be bound to AI agents in your organization; this may not be enabled for your organization yet — contact Barndoor.
The summary cards
Seven cards sit at the top, reflecting the selected window and scope.
Active Users is a peak, not a total. It reports the highest distinct-user count seen in any one period, so a window where ten users were each active on a different day shows the busiest day’s count rather than ten. Changing the granularity changes it.
Requests
One row per request that reached an outcome: success, error, or denied (blocked by a budget or rate limit, for example). Blocked requests count here even though they never reached a provider. Two kinds of traffic count as requests that you might not expect:- Routing policy decisions. When a routing policy uses a model call to pick the route, that call is recorded as a separate request under the same user and key, so one routed user request can add two to Requests. Group the Breakdown tab by Routing Role to separate them.
- Routing policy previews. Each preview an admin runs from the routing policy editor is recorded as a request under that admin, with no API key.
Tokens Used
Fresh input tokens + output tokens + cache read tokens + cache write tokens. This is the same basis the cost is computed from, so dividing Est. Cost by Tokens Used gives a sensible blended rate.- Input tokens never include cached tokens, for any provider. Cache reads and cache writes are counted separately and added on top.
- Reasoning tokens are already inside output tokens. Models that “think” before answering bill that thinking as output, and the gateway counts it there. There is no separate reasoning figure on the dashboard, and nothing to add.
- Blocked and most failed requests contribute no tokens, since no model produced a response.
Est. Cost
The sum of the cost recorded on each request when it was served, using the pricing rule that applied at that moment. It covers input, output, cache read and cache write cost — including any long-context surcharge on the matching rule. The figure is estimated because it comes from your pricing rules in Barndoor, not from the provider’s bill. If your rules match your contract, it should track the invoice closely. Some requests record $0.00:- Unpriced models — no pricing rule and no Barndoor default covers the model. The request succeeds, but no cost is recorded.
- Providers with Calculate token cost turned off — for example a provider billed on a flat subscription. Their tokens are still counted, but cost is recorded as $0 by design.
- Blocked requests, which never reach a provider.
Avg Duration
The time the gateway spent on each request, most of which is time waiting on the upstream provider. It is the whole call’s wall-clock — not the overhead Barndoor adds, and not a number to subtract from anything. Blocked requests, and requests rejected before reaching a provider (for example, an unknown model name), have no duration and are left out.The card averages each period’s average rather than every request. A quiet period counts as much as a busy one, so the card can move when you change the granularity, and one slow call in a quiet hour pulls it up. Treat it as a trend, and read the Avg Duration Over Time chart to see when it changed.
Error Rate
Requests with an error plus requests blocked by the gateway, as a share of Requests. The count beneath the percentage covers both. A spike is not necessarily a broken provider — it may be a budget or rate limit correctly denying traffic. The Errors tab separates the two: each error is attributed either to the Gateway (Barndoor’s own limits and rejections) or to the Provider.Failovers
Failed provider attempts that the gateway retried on another route, plus, beneath, the number of requests recovered that way. This is separate from Requests and Error Rate. A provider can be failing badly here while users see no errors, because failover absorbed the failures. The Resilience tab shows the detail. See Failover, Cooldowns, and Route Health.The tabs
Trends
Six charts, one value per period:- Requests Over Time
- Tokens Used Over Time — stacked into input and output, plus cache read and cache write when the window used prompt caching
- Estimated Cost Over Time — split the same way
- Avg Duration Over Time
- Active Users Over Time — distinct users in each period
- Avg Requests per User — requests ÷ active users in each period
Throughput
Peak Requests / min and Peak Tokens / min: the busiest single minute inside each period. Use them to size requests-per-minute and tokens-per-minute rate limits against real bursts, which an hourly or daily average hides.Peak Tokens / min counts each request’s total as the provider reported it. For some providers that total leaves out cached input tokens, so it isn’t on exactly the same basis as Tokens Used.
Errors
Three charts:- Error Count by Status Code — failed requests stacked by HTTP status and who issued it, for example
429 Rate limited · Gatewayversus429 Rate limited · Provider. A 429 from the gateway means one of your rate limits fired; a 429 from the provider means the provider throttled you. - Error Count by Model — failed requests stacked by the model the caller asked for.
- Error Rate (%) — errors as a share of requests in each period.
Resilience
Retry and failover activity, kept apart from the request figures on every other tab:- Provider Failures That Triggered Failover — failed attempts, stacked by the provider that failed.
- Requests Recovered by Failover vs Failed on All Routes — of the requests that hit a failure, how many succeeded on a later route and how many ran out of routes and returned an error.
- 429 Retries Absorbed Within a Route — rate-limit (429) responses the gateway retried on the same route instead of failing over or returning the error to the caller.
Caching
This tab only appears when some request in the window used prompt caching. Cache reads (billed at a discount) and cache writes (billed at a premium) are always kept separate, because they move cost in opposite directions.
Cache Savings can be negative. Writing to the cache costs more than fresh input, so a workload that writes a lot and reads little back pays more with caching than without it. The card shows ”—” when the window has no fresh input cost to compare against.
Beneath the cards, Cache Hit Rate Over Time, Cache Tokens Over Time and Cache Cost Over Time show whether reuse is improving or decaying.
By period
One row per period — hourly, daily or weekly, chosen from the Usage by dropdown — newest first, with requests, input tokens, output tokens, cache read and cache write (when used), Total tokens, Est. cost, errors and error rate. Totals for the window match the summary cards.- Quiet periods are hidden by default. The line above the table says how many are held back (“5 of 7 days with activity”); clear Hide quiet days to show them as zero rows.
- Expanding a row breaks that single period down along the dimension set in Expand to, so you can go from “Tuesday was expensive” to “Tuesday was expensive on this model” without changing the window.
- The first and last periods of a window are usually partial. A Last 7 days window starts at the current time of day seven days ago, so its first daily row covers only part of that day.
llm-usage-by-period-2026-09-01_2026-09-07-UTC.csv. Hidden quiet periods are left out of the file too.
Breakdown
The same traffic grouped along one dimension at a time, chosen from Group by, with requests, input, output and total tokens, errors, and estimated cost per row. Cache columns appear when any row used caching; a Failovers column appears for Model, Upstream Model and Provider. Switch between table, bar and pie with the toggle above it — the charts show the 15 largest rows by requests.
The metric definitions above apply unchanged. A few groupings need more care:
- Requests with no value for the dimension are left out, so the rows may not add up to the summary cards. Grouping by User, for example, omits traffic from keys not tied to a user.
- Group and Role can double count. A user in three groups contributes their full usage to each of the three rows, so group rows can add up to more than the organization total. Users with no group don’t appear at all.
- Routing Role reports all traffic. User requests is everything that isn’t a routing decision, so the rows always add up to the total. Routing overhead is what routing cost you on top of the requests it routed; Routing preview (admin) is spend from admins testing policies, which is not charged to anyone’s budget or rate limits. See Routing policies; they may not be enabled for your organization yet — contact Barndoor.
- Launch Profile only has rows if your organization runs agents through
barndoor runlaunch profiles, which may not be enabled for your organization yet — contact Barndoor. - Credential is the grouping to use when reconciling against a provider invoice. Several providers can share one credential, so a Provider breakdown can split one invoice line across several rows.
- Status Code puts a gateway 429 and a provider 429 on separate rows. The Errors column is hidden for Status Code and Error Source, where it would say nothing new.
Traffic excluded from the rows
Three groupings report requests that have no value for the dimension beneath the table rather than as a row, because that bucket is usually the largest and would crowd out the rows that matter. The note beneath the table gives its request and token count, so the rows plus the note reconcile with the summary cards.Model, Upstream Model and Provider rows are grouped by the name recorded when the request was made, so a provider renamed mid-window appears as two rows. AI Agent, Credential, API Key and Launch Profile rows are grouped by ID, so a rename keeps one row.
Per-agent LLM usage
Each AI agent’s detail page (AI Agents → select an agent) has an LLM Usage card covering the agent’s traffic over the page’s selected period, with Cost, Requests, Tokens Used and Error rate, a cost sparkline, and a By model table showing each model’s share of the agent’s cost. These use the same definitions as the dashboard: Tokens Used includes cache tokens, and usage is matched on the agent recorded on each request, so it doesn’t change when keys are unbound or rebound. The ”/ day” figures divide by the length of the whole period, including days with no traffic. When the agent has errors, the Error rate figure links to this dashboard’s Errors tab, already scoped to the agent over the same window (up to 30 days). The card only appears once API keys can be bound to AI agents in your organization; this may not be enabled for your organization yet — contact Barndoor.Frequently asked
Why doesn't Est. Cost match my provider invoice?
Why doesn't Est. Cost match my provider invoice?
Est. Cost is computed from your pricing rules at the time each request was served, not read from the provider. It drifts from the invoice when a rule’s rates don’t match your contract, when models are unpriced or a provider has Calculate token cost turned off, or when prices changed partway through the period (history is never re-costed). Check the timezone too: invoices are usually in UTC. To compare like for like, set the window to the invoice period, keep UTC, and group the Breakdown tab by Credential. See Managing Model Pricing.
Why is Tokens Used much larger than input plus output?
Why is Tokens Used much larger than input plus output?
It includes cache read and cache write tokens, which the input figure never does. Agents that reuse long prompts — coding assistants in particular — can read far more tokens from the cache than they send fresh. The Caching tab shows the split.
Why are there more requests than my users sent?
Why are there more requests than my users sent?
Requests routed through a routing policy record a second request for the policy’s own route decision, and routing policy previews are recorded too. Group the Breakdown tab by Routing Role to see how much of the total each accounts for. Blocked requests also count as requests.
My error rate spiked but the provider looks healthy.
My error rate spiked but the provider looks healthy.
Requests blocked by budgets and rate limits count as errors. On the Errors tab, check whether the spike is labelled Gateway (your own limits) or Provider. If failover is configured, the Resilience tab shows provider failures that never became user-facing errors.
Why don't the per-group rows add up to the organization total?
Why don't the per-group rows add up to the organization total?
A user who belongs to several groups is counted in each of them, and users in no group aren’t shown. Use By Group in the scope filter when you need one group’s exact totals.
Why does an expanded row show more than the row itself?
Why does an expanded row show more than the row itself?
On a rolling window the first period is partial — the row counts traffic from the start of the window, but expanding it breaks down the whole period, including the part before the window began. Pick whole days on the calendar to avoid this.
Where is MCP traffic?
Where is MCP traffic?
On the MCP Usage Dashboard (Reporting → MCP Usage Dashboard). This dashboard covers LLM Gateway traffic only.