Try Bifrost Enterprise free for 14 days. Request access

Best LLM Gateways for Monitoring Claude Code Token Spend

Claude Code token cost monitoring attributes every request's tokens and dollar cost to a developer, team, or project. This comparison covers Bifrost, LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong, plus Claude Code's native /usage and OpenTelemetry options.

Best LLM Gateways for Monitoring Claude Code Token Spend

TL;DR

  • An LLM gateway attributes Claude Code token cost to a developer, team, or project by authenticating every request with its own key before forwarding it to the model provider.
  • Claude Code's built-in /usage command and OpenTelemetry export report tokens and cost, but only a gateway can enforce a budget before a request is billed.
  • Bifrost tracks Claude Code token usage per virtual key, team, customer, and provider, and exposes input tokens, output tokens, and cost in USD as Prometheus metrics.
  • Bifrost adds 11 microseconds of overhead per request at 5,000 requests per second and routes Claude models across Anthropic, AWS Bedrock, Google Vertex, and Azure.
  • LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong AI Gateway all document Claude Code routing; they differ in where budgets are enforced and whether the gateway can be self-hosted.

A single Claude Code session can issue hundreds of model calls across Sonnet, Opus, and Haiku, and without a gateway in front of it, that Claude Code token cost lands on one provider invoice with no per-developer, per-project, or per-model breakdown. Bifrost, the open-source AI gateway built in Go by Maxim AI, is free to self-host and is the best overall choice for engineering teams that need to monitor and control Claude Code token spend across developers, projects, and models. This post compares the LLM gateways that sit between Claude Code and your providers, and explains how each one handles cost tracking, budgets, and token usage visibility. The criteria below focus on what platform teams actually need when Claude Code moves from a few pilot users to an organization-wide rollout.

Why Claude Code Token Spend Is Hard to Monitor

Claude Code token spend is hard to monitor because every request is billed to one organization account, while the questions a platform team asks are about individual developers, teams, and projects. A shared API key or a cloud provider account reports only the aggregate, and Claude Code talks directly to the Anthropic API by default.

When one developer uses it, the cost shows up on one invoice. When a hundred engineers use it concurrently across projects and approval modes, the spend becomes opaque and attribution breaks down. Anthropic's own guidance on managing Claude Code costs puts the enterprise average at around $13 per developer per active day and $150-250 per developer per month, which makes a 100-developer rollout a five-figure monthly line item.

The specific monitoring gaps that appear at scale are consistent:

  • No per-user attribution. A shared key cannot tell you which developer or team generated which tokens.
  • No budget enforcement. A runaway agent loop on a large codebase can consume thousands of dollars in tokens before anyone notices.
  • No model-level breakdown. Claude Code mixes Sonnet, Opus, and Haiku in a single session, and raw invoices rarely separate spend by tier.
  • No real-time signal. Provider billing dashboards update on a delay, so cost spikes are visible only after they have already happened.

Routing Claude Code through a gateway closes these gaps. By pointing Claude Code at a gateway base URL, every request is intercepted at the network layer and tagged with an identity before it reaches Anthropic. Bifrost supports this by reading a virtual key once the ANTHROPIC_BASE_URL environment variable points Claude Code at the gateway, which requires no change to developer workflows. The broader case for this layer is covered in our explainer on how a Claude Code gateway handles routing, governance, and cost control.

Without a gateway, developers share one API key and one invoice; with Bifrost, each developer's virtual key attributes Claude Code token usage before requests reach providers

Figure 1: A gateway turns one aggregate Claude Code bill into spend attributed to each developer's key.

How to Check Claude Code Usage Without a Gateway

Claude Code ships three native ways to check usage: the /usage command for the current session, Claude Console reporting and workspace spend limits for API organizations, and OpenTelemetry export for organization-wide metrics. They report tokens and cost well, but for API keys and cloud provider accounts none of them caps an individual developer's spend before a request is billed.

Native option What it reports Where it falls short for teams
/usage command Session tokens by model and an estimated cost at list price Local to one developer's session; totals reset with /clear
Claude Console (API organizations) Usage page, a dedicated "Claude Code" workspace, and workspace spend limits Limits apply to the workspace as a whole, not to individual developers or projects
OpenTelemetry export claude_code.token.usage and claude_code.cost.usage metrics, broken down by user, team, and model Reports after the fact; it cannot cap spend
Amazon Bedrock, Google Cloud, Microsoft Foundry Your cloud billing console and budget controls Anthropic's analytics do not cover this traffic, so per-user reporting needs OpenTelemetry or a gateway, such as Anthropic's self-hosted Claude apps gateway or an LLM gateway

Claude Code's OpenTelemetry export is enabled per machine with CLAUDE_CODE_ENABLE_TELEMETRY=1, so coverage depends on every developer's configuration. A gateway is the only option in this list that sits in the request path, so it is the only one that can reject a request once a developer's or team's budget is spent. Teams comparing gateways for this job specifically can also read our roundup of the best AI gateway to monitor Claude Code token usage.

What to Look for in an LLM Gateway for Claude Code Cost Tracking

An LLM gateway for Claude Code cost tracking is a control layer that intercepts Claude Code requests, attributes them to an identity, and records token usage and cost before forwarding the call to a model provider. When evaluating gateways for this job, weigh them against the following criteria:

  • Per-identity token tracking: can the gateway attribute token usage and cost to individual developers, teams, or projects, not just one shared key.
  • Budget controls: can it enforce hard spending limits and reject requests once a budget is exhausted.
  • Rate limiting: can it cap requests and tokens per consumer to stop runaway sessions.
  • Real-time observability: does it expose live cost and token metrics rather than delayed billing reports.
  • Multi-provider routing: can it run Claude models across Anthropic, Bedrock, Vertex, and Azure without changing developer workflows.
  • Self-hosting and data control: can it be deployed inside your own infrastructure so request data never leaves your network.

Bifrost meets all six criteria as a free, open-source layer, with governance features available in the open-source build rather than gated behind an enterprise tier. For a narrower look at the reporting side alone, see how a gateway works as a Claude Code token usage monitor for engineering teams.

The Best LLM Gateways for Claude Code Token Cost Monitoring

The best LLM gateways for Claude Code token cost monitoring are Bifrost, LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong AI Gateway. All five document a Claude Code integration; they differ in how finely they attribute spend, whether budgets are enforced before a request is forwarded, and whether the gateway runs inside your own network.

The gateways below are ranked by how completely they cover the cost-monitoring criteria above for Claude Code specifically. Competitor details reflect each vendor's public documentation as of September 2026.

Gateway Per-identity tracking Budget controls Rate limiting Real-time observability Claude routing Self-hosted
Bifrost Virtual key, team, customer, and provider level Hierarchical, enforced pre-request Requests and tokens per virtual key and provider Prometheus and OpenTelemetry Anthropic, Bedrock, Vertex, Azure Yes
LiteLLM Key, user, and team level Per key, team, and team member TPM and RPM limits Prometheus spend metrics Anthropic, Bedrock, Vertex, Azure Yes
Cloudflare AI Gateway Custom metadata or Cloudflare Access identity Spend limits by model, provider, or metadata Yes Analytics and User Insights dashboards Anthropic, Bedrock, Vertex No, managed only
OpenRouter API key, app, and user in the Activity dashboard Workspace budgets (Enterprise plan) Not published per consumer Activity dashboard Anthropic-compatible endpoint No, hosted only
Kong AI Gateway Consumer and consumer group Cost-based limits via an enterprise plugin Yes, via plugin Prometheus and OpenTelemetry metrics Anthropic, Bedrock, and others via plugin Yes, self-managed

1. Bifrost

Bifrost homepage presenting the open-source enterprise AI gateway with its npx quickstart command and throughput figures

Bifrost is a high-performance, open-source AI gateway that unifies access to 10,000+ models across 25+ providers through a single OpenAI-compatible API, and it adds only 11 microseconds of overhead per request at 5,000 requests per second in sustained benchmarks. For Claude Code, it intercepts requests at the network layer and attributes every token to a virtual key, the primary governance entity that carries its own access permissions, budgets, and rate limits. Cost is tracked independently at the customer, team, virtual key, and provider levels, so a platform team can see exactly which developer or project drove which portion of Claude Code token spend. Bifrost runs Claude models across Anthropic, AWS Bedrock, Google Vertex, and Azure, and developers can switch tiers mid-session with the /model command without touching billing.

Teams that also want to route Claude Code to non-Anthropic models can follow the guide to multi-model routing for Claude Code.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

2. LiteLLM

LiteLLM homepage describing its AI gateway for platform teams that routes humans, agents, and machines to LLMs

LiteLLM is an open-source proxy that exposes a unified API across multiple providers and tracks spend per key, user, and team. Teams routing Claude Code through it can set budgets on keys, teams, and individual team members, and export a litellm_spend_metric counter to Prometheus. LiteLLM enforces those budgets against spend stored in its database, so a deployment without a database does not cap spend, and performance overhead remains a common evaluation point for high-throughput coding workloads. Teams comparing the two can review Bifrost as a drop-in LiteLLM alternative with a full feature breakdown.

Best for: smaller teams that want a lightweight open-source proxy with key- and team-level spend tracking and do not yet need enterprise governance.

3. Cloudflare AI Gateway

Cloudflare AI Gateway product page describing a control plane for routing, usage, billing, and logs

Cloudflare AI Gateway is a hosted gateway that logs requests, caches responses, and reports analytics for traffic passing through it. It documents a Claude Code integration for Anthropic, Amazon Bedrock, and Google Vertex AI, attributes spend to users through custom metadata or Cloudflare Access identities, and can block requests with spend limits scoped by model, provider, or metadata. Cloudflare describes those spend limits as eventually consistent, so a burst of concurrent requests can briefly exceed a limit. It is a managed service rather than a self-hosted layer, which matters for teams that need request data to stay inside their own network.

Best for: teams already standardized on Cloudflare that want hosted analytics and caching with minimal setup.

4. OpenRouter

OpenRouter homepage presenting a unified interface for many models with provider and model counts

OpenRouter is a model marketplace that aggregates providers behind a single API. Claude Code connects to it by setting ANTHROPIC_BASE_URL to OpenRouter's Anthropic-compatible endpoint, and the Activity dashboard breaks down spend, requests, and tokens by model, provider, API key, app, or user. Workspace budgets that block requests at a dollar limit are available on the Enterprise plan. OpenRouter is hosted only, and every request passes through its infrastructure before reaching the model provider.

Best for: individual developers and small teams experimenting with many models through one account, rather than governed team deployments.

5. Kong AI Gateway

Kong homepage presenting AI connectivity for securing and cost-controlling LLM, MCP, and API requests

Kong AI Gateway extends the Kong API gateway with AI-specific plugins for routing and request logging. Organizations already running Kong for general API management can route Claude Code through the AI Proxy plugin and track cost with an ai_llm_cost_total metric, which requires defining input and output costs for each model. Cost-based limits per consumer or consumer group come from the AI Rate Limiting Advanced plugin, which is part of Kong's AI Gateway Enterprise offering.

Best for: infrastructure teams already invested in the Kong ecosystem for general API gateway management.

How Bifrost Tracks Claude Code Token Spend

Bifrost tracks Claude Code token spend by authenticating each request with a virtual key, checking that key's budgets and rate limits before forwarding it, and recording input tokens, output tokens, and cost in USD for every call. The records feed request logs, Prometheus metrics, and OpenTelemetry traces, so spend is visible per developer, team, model, and provider.

Bifrost records Claude Code token spend through three layers that work together: identity, metrics, and cost reduction.

The first layer is identity. Each developer or team receives a virtual key, and Claude Code sends the key set in ANTHROPIC_AUTH_TOKEN as an Authorization: Bearer header. Because every request carries a key, Bifrost attributes token usage and cost to the right consumer automatically. Budgets are checked hierarchically across customer, team, virtual key, and provider configuration levels through the hierarchical governance controls, and budgets and rate limits cap both tokens and requests per consumer to stop a runaway session before it drains a budget. The same pattern for per-team rollouts is covered in the guide to governing Claude Code token usage per team.

Budget hierarchy in Bifrost from customer to team to virtual key to provider configuration, each level with its own Claude Code spend limit

Figure 2: A Claude Code request must fit every budget above its virtual key, so a team cap holds even when individual keys still have room.

In Bifrost Enterprise, projects add a per-request accounting scope on top of the caller's identity: a developer keeps one key, and a header assigns each call's spend to a project with its own budget and its own line in every report.

The second layer is metrics. Bifrost exposes telemetry through a native Prometheus endpoint that tracks success and error rates, token usage, and real-time cost in USD, with a dedicated bifrost_cost_total metric. Teams can wire this into an existing observability stack using standard OpenTelemetry export or scrape the endpoint with Prometheus and visualize spend in Grafana. Custom labels let teams slice cost by project, environment, or any dimension they inject at request time through x-bf-dim-* headers, and alerts can fire when daily provider cost crosses a threshold.

Prometheus metric What it answers for Claude Code
bifrost_input_tokens_total How many prompt and context tokens each key, model, and provider sends
bifrost_output_tokens_total How many tokens Claude Code generates in responses and edits
bifrost_cost_total What each key, team, model, and provider costs in USD
bifrost_cache_hits_total How many requests the gateway served from cache instead of billing the provider

Every request also lands in the built-in request logs with its tokens, cost, and latency, filterable by provider, model, token range, and cost range. Enterprise alerting evaluates budget and rate-limit usage every 60 seconds and notifies Slack, Microsoft Teams, PagerDuty, or a webhook when a virtual key, team, or customer crosses a threshold.

The third layer is reducing the tokens that get billed in the first place. Semantic caching returns cached responses for identical or semantically similar requests; by default it skips conversations longer than three messages, so for Claude Code it helps most with short, repeated prompts and scripted runs rather than long agentic sessions. Claude Code also sends an x-claude-code-session-id header, which Bifrost uses for session affinity so a session keeps hitting the same provider key and its prompt cache.

For agentic, tool-heavy sessions, Code Mode lets the model write Python to orchestrate multiple tool calls in one step, which has delivered up to 92% lower token costs for MCP workloads at scale. Together these reduce the token volume that monitoring then has to account for, and the guide on how to reduce Claude Code token costs covers the full set of techniques.

For regulated teams, Bifrost Enterprise adds audit logs, RBAC, SSO, and in-VPC deployment, so Claude Code token spend can be tracked under SOC 2, HIPAA, or GDPR requirements without sending request data outside the organization.

Setting Up Claude Code Cost Monitoring with Bifrost

Routing Claude Code through Bifrost takes two required environment variables in settings.json: ANTHROPIC_BASE_URL, which points Claude Code at the gateway's Anthropic-compatible endpoint, and ANTHROPIC_AUTH_TOKEN, which carries the developer's virtual key. Two optional variables pin the Haiku and Sonnet tiers to specific models. After setting up the gateway locally or in your cluster, point Claude Code at it and supply a virtual key:

"env": {
  "ANTHROPIC_BASE_URL": "http://localhost:8080/anthropic",
  "ANTHROPIC_AUTH_TOKEN": "your-virtual-key",
  "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5",
  "ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-4-6"
}

With this configuration, every Claude Code request flows through the gateway, is authenticated by the virtual key, and is recorded against that key's budget and rate limits. No Anthropic account login is required when using the ANTHROPIC_AUTH_TOKEN method, since billing routes through the virtual key.

Claude Code request pipeline through Bifrost: virtual key authentication, budget and rate-limit check, provider call, then token and cost recording to logs and metrics

Figure 3: Budgets are checked before the provider is called, and tokens and cost are recorded after it responds.

The full setup, including provider-specific model pinning for Bedrock, Vertex, and Azure, is covered in the Claude Code integration guide, and the same pattern applies to other terminal coding agents.

For organization-wide rollouts with SSO and centrally issued keys, see the walkthrough on enterprise Claude Code management and cost control. Claude Code itself can be installed from the official Anthropic documentation.

Frequently Asked Questions

Does routing Claude Code through a gateway change the developer experience?

No. Claude Code behaves identically; the only change is the base URL and auth token in settings.json. Model switching with /model and standard workflows continue to work, while the gateway adds cost tracking transparently.

Is there a way to track Claude Code usage per developer?

Yes. Claude Code's OpenTelemetry export reports token and cost metrics per user, and a gateway adds enforcement on top. By issuing each developer a separate virtual key, Bifrost attributes token usage and cost to the individual key, and rolls those costs up to team and customer levels for reporting and budget checks.

How is Claude Code usage measured?

Claude Code usage is measured in tokens: input, output, cache read, and cache write tokens for each model. On API billing, cost is those tokens multiplied by each model's per-token rates, which is what the /usage command estimates locally. On Pro and Max subscriptions, usage counts against the plan's limits instead. A gateway such as Bifrost measures the same tokens at the request level and attaches them to a virtual key.

Is Bifrost free to use for cost monitoring?

Yes. Virtual keys, budgets, rate limits, telemetry, and cost tracking are part of the open-source build. Enterprise features such as RBAC, SSO, and in-VPC deployment are available when teams need them.

How much does Claude Code cost?

Claude Code billing depends on how it is accessed. Subscription plans bundle usage into a flat monthly fee, while API access bills per input and output token at model-specific rates, and Anthropic reports an enterprise average of about $13 per developer per active day. Only the API path produces the per-request token costs a gateway can attribute, which is why teams standardizing on API keys route Claude Code through a gateway to see spend per developer rather than one aggregate invoice.

Can Claude Code run through a proxy?

Yes. Claude Code reads its API base URL from an environment variable, so pointing it at a gateway needs no code changes and no change to how developers work. Requests then flow through the gateway, which records tokens and cost per identity before forwarding to Anthropic. The Bifrost setup docs for Claude Code cover the exact variables.

How do you reduce Claude Code token costs?

Two mechanisms do most of the work at the gateway. Semantic response caching returns a stored response when a new prompt is semantically equivalent to an earlier one, avoiding a second billed call on short, repeated prompts. Code Mode for MCP tools has the model write code to orchestrate tools instead of issuing many individual tool calls, which cuts token consumption on MCP-heavy agentic sessions.

Get Started with Bifrost

Monitoring Claude Code token cost comes down to putting an identity-aware gateway between your developers and your providers, then reading the cost and token metrics it records. Bifrost does this as a free, open-source layer with hierarchical budgets, real-time telemetry, and token-reducing features like semantic caching and Code Mode, and it scales to enterprise governance when you need it. For MCP-heavy Claude Code sessions, the breakdown of cutting Claude Code token costs with Bifrost as an MCP gateway shows where the largest savings come from.

Explore the full set of Bifrost resources or book a demo with the Bifrost team to see how LLM gateway cost tracking fits your Claude Code rollout.