---
description: Provider invoices bill per API key, not per team. Track LLM usage and spend by team with Bifrost virtual keys, hierarchical budgets, and cost labels.
title: How to Track LLM Usage and Spend by Team
image: https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/how-to-track-llm-usage-and-spend-by-team-bifrost-isometric.png
---

Try Bifrost Enterprise free for 14 days. [Request access](https://www.getmaxim.ai/#enterprise-trial)

## TL;DR

- Provider invoices group spend by API key or project, so tracking LLM usage by team requires attribution upstream of the provider.
- Bifrost virtual keys attach to one team or one customer, and every request is priced and logged with that team identity.
- Bifrost Prometheus metrics carry `team_id`, `team_name`, `customer_id`, and `customer_name` as base labels, so `bifrost_cost_total` can be grouped by team without extra headers.
- Budgets at the provider config, virtual key, team, and customer levels are checked independently, and a failure at any level blocks the request before spend occurs.
- Enterprise access profiles issue per-user virtual keys automatically, with isolated budget and rate-limit counters.

A shared OpenAI or Anthropic organization key produces one invoice line per month, which means any company running several teams against that key cannot separate LLM usage and spend by team. [Bifrost](https://www.getmaxim.ai/bifrost), the [open-source AI gateway](https://github.com/maximhq/bifrost) built in Go by Maxim AI, resolves this at the infrastructure layer: every request authenticates through a virtual key that carries team identity, and every response is priced and recorded before it returns to the caller. This post covers the full path, from labeling requests through to enforcing per-team budgets that block overspend instead of reporting it after the invoice arrives.

## Why LLM Spend Is Hard to Attribute by Team

Attributing LLM usage and spend by team is difficult because providers bill at the credential boundary rather than the organizational boundary. An API key or provider project is the finest dimension the invoice knows about, so any team structure that does not map one-to-one onto keys or projects is invisible to billing. Attribution has to be added upstream of the provider, at the point where requests are issued.

The specific obstacles teams run into:

- **Credential-level billing**: provider dashboards report totals per key or per project, not per team, feature, or customer.
- **Key splitting trade-offs**: issuing one provider key per team buys attribution but fragments failover, load balancing, and shared rate limit headroom across keys.
- **Inconsistent application tagging**: when each service adds its own metadata, coverage depends on every team remembering to instrument, and gaps appear silently.
- **Multi-provider pricing**: a team running OpenAI, Anthropic, and Bedrock needs three pricing schemas normalized into one currency figure before any rollup is meaningful.
- **Hidden cost multipliers**: retries, cached tokens, batch discounts, and reasoning tokens all change the real cost of a request in ways a raw token count does not capture.

The pressure to solve this is now widespread. The FinOps Foundation's [State of FinOps 2026 report](https://www.finops.org/update/state-of-finops-2026-report-now-available/) found that 98% of respondents manage AI spend, up from 31% two years earlier and 63% in 2025. Centralizing attribution at the gateway is one of the practical answers, and the [governance capabilities](https://www.getmaxim.ai/bifrost/resources/governance) available at that layer are what make per-team reporting possible without touching application code.

## Three Approaches to LLM Cost Attribution

Teams generally choose between provider key splitting, application-level tagging, and gateway-level attribution. The three differ in how much application code they require, how complete their coverage is, and whether they can enforce limits rather than only report on them.

| Approach | Coverage | Code changes | Can enforce limits |
| --- | --- | --- | --- |
| One provider key per team | Complete per key, but breaks cross-key routing | None | No, only the provider's own key caps |
| Application-level tagging | Only services that implement it | Every call site, every service | No, reporting only |
| Gateway-level attribution | Every request that transits the gateway | Base URL change only | Yes, at request time |

Of the three, gateway-level attribution is the approach that both covers all traffic and blocks a request before the spend happens. Because [the Bifrost AI gateway](https://www.getmaxim.ai/bifrost) works as a drop-in replacement for existing SDKs, adopting it means changing a base URL rather than rewriting call sites, which is what makes complete coverage realistic. For teams comparing options at this layer, the [LLM Gateway Buyer's Guide](https://www.getmaxim.ai/bifrost/resources/buyers-guide) breaks down the capability matrix in more detail, and this roundup of [LLM cost tracking tools](https://www.getmaxim.ai/articles/best-llm-cost-tracking-tools-in-2026/) compares dedicated tracking products with gateway-based approaches.

## How Bifrost Tracks LLM Usage and Spend by Team

Virtual keys are the primary governance entity in [the open-source Bifrost gateway](https://www.getmaxim.ai/bifrost). Applications authenticate with a virtual key instead of a raw provider key, and that key carries the access permissions, budget, and rate limits for whoever holds it. Requests can present the key through the `x-bf-vk` header or through the standard `Authorization`, `x-api-key`, `x-goog-api-key`, and `api-key` headers, so existing SDK code continues to work unchanged.

Each [virtual key](https://docs.getbifrost.ai/features/governance/virtual-keys) attaches to exactly one team or one customer, and those attachments are mutually exclusive. That constraint is what makes team rollups unambiguous: a request's spend belongs to one team, never two. A virtual key can also attach directly to a customer or stand alone, so the team level is optional rather than required.

The [budget hierarchy](https://docs.getbifrost.ai/features/governance/budget-and-limits) runs up to four levels deep when a virtual key belongs to a team, with an independent budget at each level:

- **Customer**: the top-level entity, typically a business unit or an external tenant
- **Team**: one-to-many under a customer, holding the team's own budget
- **Virtual key**: one-to-many under a team, with its own budget and rate limits
- **Provider config**: one-to-many under a virtual key, with per-provider budgets and rate limits

Cost is computed by Bifrost rather than estimated. Pricing is applied using model pricing data from the [Model Catalog](https://docs.getbifrost.ai/architecture/framework/model-catalog), which syncs automatically every 24 hours by default, along with actual input and output token counts from the provider response, the request type (chat, text, embedding, speech, or transcription), cache status for reduced-cost cached responses, and batch operation discounts. The result is a per-request dollar figure recorded alongside the rest of the request metadata for that call. The [request logs](https://docs.getbifrost.ai/features/observability/default) can be filtered by cost range, provider, model, token usage, and time range, which gives finance and platform teams a per-request audit trail behind every team total.

## Labeling Requests for Per-Team Reporting

Per-team reporting in Bifrost starts from labels the gateway already attaches. When governance is in use, every upstream metric carries the virtual key, team, and customer as base labels (`virtual_key_id`, `virtual_key_name`, `team_id`, `team_name`, `customer_id`, `customer_name`), so a request made with a team-attached virtual key is attributed to that team with no extra instrumentation.

For dimensions the governance hierarchy does not model, such as environment, project, or feature, [Bifrost, the AI gateway](https://www.getmaxim.ai/bifrost), accepts `x-bf-dim-*` headers and turns each one into a label on the metrics it emits, so a project name travels with the request from the caller to the dashboard.

Configure the label set once, then send values per request:

```bash
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-bf-dim-team: engineering" \
  -H "x-bf-dim-environment: production" \
  -H "x-bf-dim-project: support-agent" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
```

The label names are declared in configuration as `prometheus_labels`, for example `["team", "environment", "organization", "project"]`. From there, [telemetry](https://docs.getbifrost.ai/features/telemetry) exposes the counters that matter for cost reporting, each carrying the provider, model, virtual key, team, and customer base labels plus the custom dimensions:

- `bifrost_cost_total`: total cost in USD for upstream provider requests
- `bifrost_input_tokens_total` and `bifrost_output_tokens_total`: token volume sent and received
- `bifrost_upstream_requests_total`: request counts, with success and error variants
- `bifrost_cache_hits_total`: direct and semantic cache hits, broken out by cache type

A single Prometheus query over `bifrost_cost_total` grouped by team produces LLM usage and spend per team without any additional pipeline:

```promql
sum by (team_name) (increase(bifrost_cost_total[30d]))
```

Grouping by the custom `team` label instead works the same way for traffic that does not use team-attached virtual keys. Metrics are available by [scraping the `/metrics` endpoint or pushing to a Push Gateway](https://docs.getbifrost.ai/features/observability/prometheus) for multi-node deployments, and telemetry collection runs asynchronously so it adds no latency to request processing. The same `x-bf-dim-*` values also propagate into internal logs and OpenTelemetry span attributes, which keeps team attribution consistent with the [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai) already used elsewhere in the stack. For a deeper look at building dashboards on top of these signals, see [per-team cost attribution as a reporting layer for AI usage](https://www.getmaxim.ai/articles/per-team-cost-attribution-a-reporting-layer-for-ai-usage/).

## Enforcing Team Budgets, Not Just Reporting Them

Reporting explains overspend after it happens; budgets prevent it. On [the Bifrost platform](https://www.getmaxim.ai/bifrost), every applicable budget in the hierarchy is checked independently before a request proceeds, and a failure at any single level blocks it.

The checking sequence for a team-attached virtual key runs provider config budget, then virtual key budget, then team budget, then customer budget. When the request succeeds, its cost is deducted from every applicable level, so a team budget and the customer budget above it both reflect the same spend.

Practical details worth configuring deliberately:

- **Reset durations**: budgets and rate limits accept durations such as `1m`, `5m`, `1h`, `1d`, `1w`, `1M`, and `1Y`. Budgets also accept `1Q` for quarterly windows, and `1d`, `1w`, `1M`, `1Q`, and `1Y` are the typical choices for cost control.
- **Calendar alignment**: setting `calendar_aligned` resets a budget at the start of each calendar period in UTC instead of on a rolling window, which lines team budgets up with finance reporting periods. Quarterly budgets can follow a fiscal calendar that starts in any month.
- **Rate limits**: [request and token limits](https://docs.getbifrost.ai/features/governance/budget-and-limits#rate-limiting) apply at the virtual key and provider config levels, running in parallel so a team can be capped on call volume and token throughput independently.
- **Routing interaction**: a provider that exceeds its budget or rate limits is excluded from routing, so other providers under the same virtual key stay available instead of the request failing outright.

This combination is what separates cost governance from cost reporting, and the [Bifrost governance resources](https://www.getmaxim.ai/bifrost/resources/governance) cover how access control, budgets, and limits compose into a single policy per consumer. The configuration details for each level are covered in [LLM budget management with virtual keys and hierarchical spend controls](https://www.getmaxim.ai/articles/llm-budget-management-virtual-keys-and-hierarchical-spend-controls/).

## Scaling Team Attribution Across a Large Organization

Manual virtual key issuance stops scaling once the number of teams passes a few dozen. [The Bifrost gateway](https://www.getmaxim.ai/bifrost) addresses this in the enterprise tier with access profiles: reusable policy templates that describe what a user, team, or business unit is permitted to do, and that materialize into virtual keys automatically.

[Access profiles](https://docs.getbifrost.ai/enterprise/access-profiles) carry the provider list, model whitelist, budgets, rate limits, and MCP tool access defined once on the template. Assigning a profile creates a per-user copy with isolated budget and rate limit counters, and marking a profile as a role's default provisions new users in that role automatically. Auto-issued keys are write-protected, so a user cannot weaken their own policy by editing the key, and every template change is recorded with a full snapshot history. With [advanced governance](https://docs.getbifrost.ai/enterprise/advanced-governance), identity provider attributes such as department or group can map users to teams, business units, and access profiles, so team attribution follows the org chart without manual key assignment.

Two further capabilities matter for organizations reporting spend across business units. Customer-scoped requests let a team attached to multiple customers attribute a single request to one specific customer using the `x-bf-customer-id` or `x-bf-customer-name` header, rather than charging every customer the team belongs to. And per-request metadata (timestamps, provider, model, latency, token counts, cost, and status) is retained in the logs store, with large payloads optionally offloaded to S3 or GCS through [log exports](https://docs.getbifrost.ai/enterprise/log-exports) so long retention stays affordable. Regulated teams running this in isolated environments can review the deployment and compliance options on the [Bifrost Enterprise](https://www.getmaxim.ai/bifrost/enterprise) page.

## Frequently Asked Questions About Tracking LLM Spend by Team

### How do you track LLM usage by team without changing application code?

Route traffic through an [AI gateway](https://www.getmaxim.ai/articles/top-5-llm-gateways-in-2026-a-production-ready-comparison/) and issue each team its own virtual key. With Bifrost, applications change only the base URL and the key they send, and existing OpenAI, Anthropic, and Gemini SDKs keep working. The gateway prices each request and records the team from the virtual key, so usage and spend by team appear in logs and metrics without instrumenting individual services.

### Can one provider API key be shared across teams while still tracking spend separately?

Yes. Bifrost holds the provider keys and teams authenticate with virtual keys instead. Every team shares the same provider credentials, failover configuration, and load balancing, while cost is recorded per virtual key and rolled up to the team and customer. This avoids the fragmentation that comes from issuing one provider key per team.

### How is the cost of each LLM request calculated?

Bifrost multiplies the input and output token counts returned by the provider by pricing data from its Model Catalog, which syncs every 24 hours by default. The calculation accounts for request type, cached responses, and batch discounts, and the resulting USD figure is written to the request log and added to the `bifrost_cost_total` metric.

### What happens when a team exceeds its LLM budget?

The request is blocked before it reaches the provider. Bifrost checks the provider config, virtual key, team, and customer budgets independently, and any single failure rejects the call. If only one provider config under a virtual key is over its budget or rate limit, that provider is excluded from routing and other providers under the same key keep serving requests.

### Can budgets reset monthly or quarterly to match finance reporting?

Yes. Setting `calendar_aligned` makes budgets reset at the start of each UTC calendar period instead of on a rolling window. Monthly budgets reset on the first day of the month, and quarterly `1Q` budgets can follow a fiscal year that starts in any month. Turning alignment on for an existing budget keeps its accumulated usage.

### How do coding agents fit into per-team LLM cost tracking?

Coding agents such as Claude Code can point at Bifrost like any other client, so their usage is priced and attributed through the same virtual keys and budgets. Giving each engineering team a virtual key for agent traffic separates agent spend from application spend, as covered in this guide to [cost tracking Claude Code with Bifrost](https://www.getmaxim.ai/articles/cost-tracking-claude-code-with-bifrost-ai-gateway/).

## Start Tracking LLM Spend by Team with Bifrost

Tracking LLM usage and spend by team comes down to three decisions: authenticate every application through a virtual key rather than a raw provider key, attach team dimensions to each request so reporting can group by them, and set budgets at the team level so overspend is blocked at request time. [The Bifrost gateway layer](https://www.getmaxim.ai/bifrost) provides all three without application rewrites, since adopting it changes a base URL rather than call sites. Teams evaluating options side by side can compare [enterprise gateways for LLM cost tracking and budget controls](https://www.getmaxim.ai/articles/top-5-enterprise-gateways-for-llm-cost-tracking-and-budget-controls/) or revisit the broader [LLM cost tracking tools landscape](https://www.getmaxim.ai/articles/best-llm-cost-tracking-tools-in-2026/).

To see how per-team budgets, cost labels, and access profiles would map onto your organization, [book a demo](https://getmaxim.ai/bifrost/book-a-demo) with the Bifrost team.

## Read next

[![Top 5 AI Gateways for Controlling Shadow AI in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-controlling-shadow-ai-bifrost-isometric.png) Shadow AI is the use of AI tools, models, and MCP servers that security teams have not approved and cannot see. This guide ranks five AI gateways for controlling it, including Bifrost with Bifrost Edge, Kong AI Gateway, Cloudflare AI Gateway, and Gravitee.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-controlling-shadow-ai/)

[![Top 5 AI Gateways for SSO and RBAC in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-sso-and-rbac-in-2026-bifrost-isometric.png) AI gateways with SSO and RBAC let enterprises tie every model request and every configuration change to a corporate identity. This guide compares Bifrost, Kong AI Gateway, Azure API Management, Gravitee, and Cloudflare AI Gateway on identity, roles, provisioning, and audit.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-sso-and-rbac-in-2026/)

[![Semantic Caching: The Top 5 AI Gateways in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/semantic-caching-the-top-5-ai-gateways-in-2026-bifrost-isometric.png) Semantic caching serves a stored LLM response when a new prompt means the same thing as an earlier one. This guide compares Bifrost, Kong AI Gateway, Azure API Management, and Cloudflare AI Gateway on match modes, vector stores, thresholds, TTLs, and cache scoping.](https://www.getmaxim.ai/articles/semantic-caching-the-top-5-ai-gateways-in-2026/)

```json
{
    "@context": "https://schema.org",
    "@type": "Article",
    "publisher": {
        "@type": "Organization",
        "name": "Maxim Articles",
        "url": "https://www.getmaxim.ai/articles/",
        "logo": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w256h256/2025/08/thumbnail.png",
            "width": 60,
            "height": 60
        }
    },
    "author": {
        "@type": "Person",
        "name": "Kamya Shah",
        "image": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2025/09/WhatsApp-Image-2025-08-29-at-17.40.40-1.jpeg",
            "width": 1200,
            "height": 1600
        },
        "url": "https://www.getmaxim.ai/articles/author/kamya/",
        "sameAs": []
    },
    "headline": "How to Track LLM Usage and Spend by Team",
    "url": "https://www.getmaxim.ai/articles/how-to-track-llm-usage-and-spend-by-team/",
    "datePublished": "2026-09-04T15:24:00.000Z",
    "dateModified": "2026-10-08T15:41:28.000Z",
    "image": {
        "@type": "ImageObject",
        "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/how-to-track-llm-usage-and-spend-by-team-bifrost-isometric.png",
        "width": 1200,
        "height": 675
    },
    "keywords": "AI Gateway",
    "description": "TL;DR\n\n * Provider invoices group spend by API key or project, so tracking LLM usage by team requires attribution upstream of the provider.\n * Bifrost virtual keys attach to one team or one customer, and every request is priced and logged with that team identity.\n * Bifrost Prometheus metrics carry team_id, team_name, customer_id, and customer_name as base labels, so bifrost_cost_total can be grouped by team without extra headers.\n * Budgets at the provider config, virtual key, team, and custome",
    "mainEntityOfPage": "https://www.getmaxim.ai/articles/how-to-track-llm-usage-and-spend-by-team/"
}
```
