---
description: Compare LLM gateways for Claude Code token cost monitoring, with per-developer attribution, pre-request budgets, and Prometheus metrics for platform teams.
title: Best LLM Gateways for Monitoring Claude Code Token Spend
image: https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/best-llm-gateways-for-monitoring-claude-code-token-spend-bifrost-weave.optimized.png
---

Try Bifrost Enterprise free for 14 days. [Request access](https://www.getmaxim.ai/#enterprise-trial)

**TL;DR**

- An [LLM gateway](https://www.getmaxim.ai/llm-gateway) attributes Claude Code token cost to a developer, team, or project by authenticating every request with its own key before forwarding it to the model provider.
- Claude Code's built-in `/usage` command and OpenTelemetry export report tokens and cost, but only a gateway can enforce a budget before a request is billed.
- Bifrost tracks Claude Code token usage per virtual key, team, customer, and provider, and exposes input tokens, output tokens, and cost in USD as Prometheus metrics.
- Bifrost adds 11 microseconds of overhead per request at 5,000 requests per second and routes Claude models across Anthropic, AWS Bedrock, Google Vertex, and Azure.
- LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong AI Gateway all document Claude Code routing; they differ in where budgets are enforced and whether the gateway can be self-hosted.

A single Claude Code session can issue hundreds of model calls across Sonnet, Opus, and Haiku, and without a gateway in front of it, that Claude Code token cost lands on one provider invoice with no per-developer, per-project, or per-model breakdown. [Bifrost](https://www.getmaxim.ai/bifrost), the [open-source AI gateway](https://github.com/maximhq/bifrost) built in Go by Maxim AI, is free to self-host and is the best overall choice for engineering teams that need to monitor and control Claude Code token spend across developers, projects, and models. This post compares the LLM gateways that sit between Claude Code and your providers, and explains how each one handles cost tracking, budgets, and token usage visibility. The criteria below focus on what platform teams actually need when Claude Code moves from a few pilot users to an organization-wide rollout.

## Why Claude Code Token Spend Is Hard to Monitor

Claude Code token spend is hard to monitor because every request is billed to one organization account, while the questions a platform team asks are about individual developers, teams, and projects. A shared API key or a cloud provider account reports only the aggregate, and Claude Code talks directly to the Anthropic API by default.

When one developer uses it, the cost shows up on one invoice. When a hundred engineers use it concurrently across projects and approval modes, the spend becomes opaque and attribution breaks down. Anthropic's own [guidance on managing Claude Code costs](https://code.claude.com/docs/en/costs) puts the enterprise average at around $13 per developer per active day and $150-250 per developer per month, which makes a 100-developer rollout a five-figure monthly line item.

The specific monitoring gaps that appear at scale are consistent:

- **No per-user attribution.** A shared key cannot tell you which developer or team generated which tokens.
- **No budget enforcement.** A runaway agent loop on a large codebase can consume thousands of dollars in tokens before anyone notices.
- **No model-level breakdown.** Claude Code mixes Sonnet, Opus, and Haiku in a single session, and raw invoices rarely separate spend by tier.
- **No real-time signal.** Provider billing dashboards update on a delay, so cost spikes are visible only after they have already happened.

Routing Claude Code through a gateway closes these gaps. By pointing Claude Code at a gateway base URL, every request is intercepted at the network layer and tagged with an identity before it reaches Anthropic. Bifrost supports this by reading a virtual key once the `ANTHROPIC_BASE_URL` environment variable points Claude Code at the gateway, which requires no change to developer workflows. The broader case for this layer is covered in our explainer on [how a Claude Code gateway handles routing, governance, and cost control](https://www.getmaxim.ai/articles/claude-code-gateway-explained-routing-governance-and-cost-control/).

![Without a gateway, developers share one API key and one invoice; with Bifrost, each developer's virtual key attributes Claude Code token usage before requests reach providers](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/best-llm-gateways-for-monitoring-claude-code-token-spend/claude-code-spend-without-vs-with-gateway.png)

*Figure 1: A gateway turns one aggregate Claude Code bill into spend attributed to each developer's key.*

## How to Check Claude Code Usage Without a Gateway

Claude Code ships three native ways to check usage: the `/usage` command for the current session, Claude Console reporting and workspace spend limits for API organizations, and OpenTelemetry export for organization-wide metrics. They report tokens and cost well, but for API keys and cloud provider accounts none of them caps an individual developer's spend before a request is billed.

| Native option | What it reports | Where it falls short for teams |
| --- | --- | --- |
| `/usage` command | Session tokens by model and an estimated cost at list price | Local to one developer's session; totals reset with `/clear` |
| Claude Console (API organizations) | Usage page, a dedicated "Claude Code" workspace, and workspace spend limits | Limits apply to the workspace as a whole, not to individual developers or projects |
| OpenTelemetry export | `claude_code.token.usage` and `claude_code.cost.usage` metrics, broken down by user, team, and model | Reports after the fact; it cannot cap spend |
| Amazon Bedrock, Google Cloud, Microsoft Foundry | Your cloud billing console and budget controls | Anthropic's analytics do not cover this traffic, so per-user reporting needs OpenTelemetry or a gateway, such as Anthropic's self-hosted Claude apps gateway or an LLM gateway |

Claude Code's OpenTelemetry export is enabled per machine with `CLAUDE_CODE_ENABLE_TELEMETRY=1`, so coverage depends on every developer's configuration. A gateway is the only option in this list that sits in the request path, so it is the only one that can reject a request once a developer's or team's budget is spent. Teams comparing gateways for this job specifically can also read our roundup of the [best AI gateway to monitor Claude Code token usage](https://www.getmaxim.ai/articles/best-ai-gateway-to-monitor-claude-code-token-usage/).

## What to Look for in an LLM Gateway for Claude Code Cost Tracking

An LLM gateway for Claude Code cost tracking is a control layer that intercepts Claude Code requests, attributes them to an identity, and records token usage and cost before forwarding the call to a model provider. When evaluating gateways for this job, weigh them against the following criteria:

- **Per-identity token tracking:** can the gateway attribute token usage and cost to individual developers, teams, or projects, not just one shared key.
- **Budget controls:** can it enforce hard spending limits and reject requests once a budget is exhausted.
- **Rate limiting:** can it cap requests and tokens per consumer to stop runaway sessions.
- **Real-time observability:** does it expose live cost and token metrics rather than delayed billing reports.
- **Multi-provider routing:** can it run Claude models across Anthropic, Bedrock, Vertex, and Azure without changing developer workflows.
- **Self-hosting and data control:** can it be deployed inside your own infrastructure so request data never leaves your network.

[Bifrost](https://www.getmaxim.ai/bifrost) meets all six criteria as a free, open-source layer, with [governance features](https://www.getmaxim.ai/bifrost/resources/governance) available in the open-source build rather than gated behind an enterprise tier. For a narrower look at the reporting side alone, see how a gateway works as a [Claude Code token usage monitor](https://www.getmaxim.ai/articles/best-ai-gateway-to-monitor-claude-code-token-usage/) for engineering teams.

## The Best LLM Gateways for Claude Code Token Cost Monitoring

The best LLM gateways for Claude Code token cost monitoring are Bifrost, LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong AI Gateway. All five document a Claude Code integration; they differ in how finely they attribute spend, whether budgets are enforced before a request is forwarded, and whether the gateway runs inside your own network.

The gateways below are ranked by how completely they cover the cost-monitoring criteria above for Claude Code specifically. Competitor details reflect each vendor's public documentation as of September 2026.

| Gateway | Per-identity tracking | Budget controls | Rate limiting | Real-time observability | Claude routing | Self-hosted |
| --- | --- | --- | --- | --- | --- | --- |
| [**Bifrost**](https://www.getmaxim.ai/bifrost) | Virtual key, team, customer, and provider level | Hierarchical, enforced pre-request | Requests and tokens per virtual key and provider | Prometheus and OpenTelemetry | Anthropic, Bedrock, Vertex, Azure | Yes |
| **LiteLLM** | Key, user, and team level | Per key, team, and team member | TPM and RPM limits | Prometheus spend metrics | Anthropic, Bedrock, Vertex, Azure | Yes |
| **Cloudflare AI Gateway** | Custom metadata or Cloudflare Access identity | Spend limits by model, provider, or metadata | Yes | Analytics and User Insights dashboards | Anthropic, Bedrock, Vertex | No, managed only |
| **OpenRouter** | API key, app, and user in the Activity dashboard | Workspace budgets (Enterprise plan) | Not published per consumer | Activity dashboard | Anthropic-compatible endpoint | No, hosted only |
| **Kong AI Gateway** | Consumer and consumer group | Cost-based limits via an enterprise plugin | Yes, via plugin | Prometheus and OpenTelemetry metrics | Anthropic, Bedrock, and others via plugin | Yes, self-managed |

### 1. Bifrost

![Bifrost homepage presenting the open-source enterprise AI gateway with its npx quickstart command and throughput figures](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2026/08/image-16.png)

Bifrost is a high-performance, open-source [AI gateway](https://www.getmaxim.ai/articles/top-5-llm-gateways-in-2026-a-production-ready-comparison/) that unifies access to 10,000+ models across 25+ providers through a single OpenAI-compatible API, and it adds only [11 microseconds of overhead per request at 5,000 requests per second](https://www.getmaxim.ai/bifrost/resources/benchmarks) in sustained benchmarks. For Claude Code, it intercepts requests at the network layer and attributes every token to a [virtual key](https://docs.getbifrost.ai/features/governance/virtual-keys), the primary governance entity that carries its own access permissions, budgets, and rate limits. Cost is tracked independently at the customer, team, virtual key, and provider levels, so a platform team can see exactly which developer or project drove which portion of Claude Code token spend. Bifrost runs Claude models across [Anthropic, AWS Bedrock, Google Vertex, and Azure](https://docs.getbifrost.ai/providers/supported-providers/overview), and developers can switch tiers mid-session with the `/model` command without touching billing.

Teams that also want to route Claude Code to non-Anthropic models can follow the guide to [multi-model routing for Claude Code](https://www.getmaxim.ai/articles/best-llm-gateways-for-claude-code-multi-model-routing/).

**Best for:** Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, [MCP gateway](https://www.getmaxim.ai/mcp-gateway), and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

### 2. LiteLLM

![LiteLLM homepage describing its AI gateway for platform teams that routes humans, agents, and machines to LLMs](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2026/08/image-17.png)

LiteLLM is an open-source proxy that exposes a unified API across multiple providers and tracks spend per key, user, and team. Teams routing Claude Code through it can set budgets on keys, teams, and individual team members, and export a `litellm_spend_metric` counter to Prometheus. LiteLLM enforces those budgets against spend stored in its database, so a deployment without a database does not cap spend, and performance overhead remains a common evaluation point for high-throughput coding workloads. Teams comparing the two can review Bifrost as a [drop-in LiteLLM alternative](https://www.getmaxim.ai/bifrost/alternatives/litellm-alternatives) with a full feature breakdown.

**Best for:** smaller teams that want a lightweight open-source proxy with key- and team-level spend tracking and do not yet need enterprise governance.

### 3. Cloudflare AI Gateway

![Cloudflare AI Gateway product page describing a control plane for routing, usage, billing, and logs](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2026/08/image-18.png)

Cloudflare AI Gateway is a hosted gateway that logs requests, caches responses, and reports analytics for traffic passing through it. It documents a Claude Code integration for Anthropic, Amazon Bedrock, and Google Vertex AI, attributes spend to users through custom metadata or Cloudflare Access identities, and can block requests with spend limits scoped by model, provider, or metadata. Cloudflare describes those spend limits as eventually consistent, so a burst of concurrent requests can briefly exceed a limit. It is a managed service rather than a self-hosted layer, which matters for teams that need request data to stay inside their own network.

**Best for:** teams already standardized on Cloudflare that want hosted analytics and caching with minimal setup.

### 4. OpenRouter

![OpenRouter homepage presenting a unified interface for many models with provider and model counts](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2026/08/image-19.png)

OpenRouter is a model marketplace that aggregates providers behind a single API. Claude Code connects to it by setting `ANTHROPIC_BASE_URL` to OpenRouter's Anthropic-compatible endpoint, and the Activity dashboard breaks down spend, requests, and tokens by model, provider, API key, app, or user. Workspace budgets that block requests at a dollar limit are available on the Enterprise plan. OpenRouter is hosted only, and every request passes through its infrastructure before reaching the model provider.

**Best for:** individual developers and small teams experimenting with many models through one account, rather than governed team deployments.

### 5. Kong AI Gateway

![Kong homepage presenting AI connectivity for securing and cost-controlling LLM, MCP, and API requests](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2026/08/image-20.png)

Kong AI Gateway extends the Kong API gateway with AI-specific plugins for routing and request logging. Organizations already running Kong for general API management can route Claude Code through the AI Proxy plugin and track cost with an `ai_llm_cost_total` metric, which requires defining input and output costs for each model. Cost-based limits per consumer or consumer group come from the AI Rate Limiting Advanced plugin, which is part of Kong's AI Gateway Enterprise offering.

**Best for:** infrastructure teams already invested in the Kong ecosystem for general API gateway management.

## How Bifrost Tracks Claude Code Token Spend

Bifrost tracks Claude Code token spend by authenticating each request with a virtual key, checking that key's budgets and rate limits before forwarding it, and recording input tokens, output tokens, and cost in USD for every call. The records feed request logs, Prometheus metrics, and OpenTelemetry traces, so spend is visible per developer, team, model, and provider.

[Bifrost](https://www.getmaxim.ai/bifrost) records Claude Code token spend through three layers that work together: identity, metrics, and cost reduction.

The first layer is identity. Each developer or team receives a virtual key, and Claude Code sends the key set in `ANTHROPIC_AUTH_TOKEN` as an `Authorization: Bearer` header. Because every request carries a key, Bifrost attributes token usage and cost to the right consumer automatically. Budgets are checked hierarchically across customer, team, virtual key, and provider configuration levels through the [hierarchical governance controls](https://www.getmaxim.ai/bifrost/resources/governance), and [budgets and rate limits](https://docs.getbifrost.ai/features/governance/budget-and-limits) cap both tokens and requests per consumer to stop a runaway session before it drains a budget. The same pattern for per-team rollouts is covered in the guide to [governing Claude Code token usage per team](https://www.getmaxim.ai/articles/best-claude-code-gateway-to-govern-token-usage-per-team/).

![Budget hierarchy in Bifrost from customer to team to virtual key to provider configuration, each level with its own Claude Code spend limit](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/best-llm-gateways-for-monitoring-claude-code-token-spend/claude-code-budget-hierarchy.png)

*Figure 2: A Claude Code request must fit every budget above its virtual key, so a team cap holds even when individual keys still have room.*

In Bifrost Enterprise, [projects](https://docs.getbifrost.ai/enterprise/projects) add a per-request accounting scope on top of the caller's identity: a developer keeps one key, and a header assigns each call's spend to a project with its own budget and its own line in every report.

The second layer is metrics. Bifrost exposes [telemetry](https://docs.getbifrost.ai/features/telemetry) through a native Prometheus endpoint that tracks success and error rates, token usage, and real-time cost in USD, with a dedicated `bifrost_cost_total` metric. Teams can wire this into an existing observability stack using standard [OpenTelemetry](https://opentelemetry.io/) export or scrape the endpoint with [Prometheus](https://prometheus.io/) and visualize spend in Grafana. Custom labels let teams slice cost by project, environment, or any dimension they inject at request time through `x-bf-dim-*` headers, and alerts can fire when daily provider cost crosses a threshold.

| Prometheus metric | What it answers for Claude Code |
| --- | --- |
| `bifrost_input_tokens_total` | How many prompt and context tokens each key, model, and provider sends |
| `bifrost_output_tokens_total` | How many tokens Claude Code generates in responses and edits |
| `bifrost_cost_total` | What each key, team, model, and provider costs in USD |
| `bifrost_cache_hits_total` | How many requests the gateway served from cache instead of billing the provider |

Every request also lands in the [built-in request logs](https://docs.getbifrost.ai/features/observability/default) with its tokens, cost, and latency, filterable by provider, model, token range, and cost range. [Enterprise alerting](https://docs.getbifrost.ai/enterprise/alerting/overview) evaluates budget and rate-limit usage every 60 seconds and notifies Slack, Microsoft Teams, PagerDuty, or a webhook when a virtual key, team, or customer crosses a threshold.

The third layer is reducing the tokens that get billed in the first place. [Semantic caching](https://docs.getbifrost.ai/features/semantic-caching) returns cached responses for identical or semantically similar requests; by default it skips conversations longer than three messages, so for Claude Code it helps most with short, repeated prompts and scripted runs rather than long agentic sessions. Claude Code also sends an `x-claude-code-session-id` header, which Bifrost uses for [session affinity](https://docs.getbifrost.ai/providers/session-affinity) so a session keeps hitting the same provider key and its prompt cache.

For agentic, tool-heavy sessions, [Code Mode](https://docs.getbifrost.ai/mcp/code-mode) lets the model write Python to orchestrate multiple tool calls in one step, which has delivered [up to 92% lower token costs for MCP workloads at scale](https://www.getmaxim.ai/bifrost/blog/bifrost-mcp-gateway-access-control-cost-governance-and-92-lower-token-costs-at-scale). Together these reduce the token volume that monitoring then has to account for, and the guide on [how to reduce Claude Code token costs](https://www.getmaxim.ai/articles/how-to-reduce-claude-code-token-costs/) covers the full set of techniques.

For regulated teams, [Bifrost Enterprise](https://www.getmaxim.ai/bifrost/enterprise) adds audit logs, RBAC, SSO, and in-VPC deployment, so Claude Code token spend can be tracked under SOC 2, HIPAA, or GDPR requirements without sending request data outside the organization.

## Setting Up Claude Code Cost Monitoring with Bifrost

Routing Claude Code through Bifrost takes two required environment variables in `settings.json`: `ANTHROPIC_BASE_URL`, which points Claude Code at the gateway's Anthropic-compatible endpoint, and `ANTHROPIC_AUTH_TOKEN`, which carries the developer's virtual key. Two optional variables pin the Haiku and Sonnet tiers to specific models. After [setting up the gateway](https://docs.getbifrost.ai/quickstart/gateway/setting-up) locally or in your cluster, point Claude Code at it and supply a virtual key:

```
"env": {
  "ANTHROPIC_BASE_URL": "http://localhost:8080/anthropic",
  "ANTHROPIC_AUTH_TOKEN": "your-virtual-key",
  "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5",
  "ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-4-6"
}
```

With this configuration, every Claude Code request flows through the gateway, is authenticated by the virtual key, and is recorded against that key's budget and rate limits. No Anthropic account login is required when using the `ANTHROPIC_AUTH_TOKEN` method, since billing routes through the virtual key.

![Claude Code request pipeline through Bifrost: virtual key authentication, budget and rate-limit check, provider call, then token and cost recording to logs and metrics](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/best-llm-gateways-for-monitoring-claude-code-token-spend/claude-code-cost-monitoring-request-path.png)

*Figure 3: Budgets are checked before the provider is called, and tokens and cost are recorded after it responds.*

The full setup, including provider-specific model pinning for Bedrock, Vertex, and Azure, is covered in the [Claude Code integration guide](https://docs.getbifrost.ai/cli-agents/claude-code), and the same pattern applies to other [terminal coding agents](https://docs.getbifrost.ai/cli-agents/overview).

For organization-wide rollouts with SSO and centrally issued keys, see the walkthrough on [enterprise Claude Code management and cost control](https://www.getmaxim.ai/articles/best-ai-gateway-for-enterprise-claude-code-management-governance-cost-control-and-monitoring/). Claude Code itself can be installed from the [official Anthropic documentation](https://code.claude.com/docs/en/overview).

## Frequently Asked Questions

### Does routing Claude Code through a gateway change the developer experience?

No. Claude Code behaves identically; the only change is the base URL and auth token in `settings.json`. Model switching with `/model` and standard workflows continue to work, while the gateway adds cost tracking transparently.

### Is there a way to track Claude Code usage per developer?

Yes. Claude Code's OpenTelemetry export reports token and cost metrics per user, and a gateway adds enforcement on top. By issuing each developer a separate virtual key, Bifrost attributes token usage and cost to the individual key, and rolls those costs up to team and customer levels for reporting and budget checks.

### How is Claude Code usage measured?

Claude Code usage is measured in tokens: input, output, cache read, and cache write tokens for each model. On API billing, cost is those tokens multiplied by each model's per-token rates, which is what the `/usage` command estimates locally. On Pro and Max subscriptions, usage counts against the plan's limits instead. A gateway such as Bifrost measures the same tokens at the request level and attaches them to a virtual key.

### Is Bifrost free to use for cost monitoring?

Yes. Virtual keys, budgets, rate limits, telemetry, and cost tracking are part of the open-source build. Enterprise features such as RBAC, SSO, and in-VPC deployment are available when teams need them.

### How much does Claude Code cost?

Claude Code billing depends on how it is accessed. Subscription plans bundle usage into a flat monthly fee, while API access bills per input and output token at model-specific rates, and Anthropic reports an enterprise average of about $13 per developer per active day. Only the API path produces the per-request token costs a gateway can attribute, which is why teams standardizing on API keys route Claude Code through a gateway to see spend per developer rather than one aggregate invoice.

### Can Claude Code run through a proxy?

Yes. Claude Code reads its API base URL from an environment variable, so pointing it at a gateway needs no code changes and no change to how developers work. Requests then flow through the gateway, which records tokens and cost per identity before forwarding to Anthropic. The [Bifrost setup docs for Claude Code](https://docs.getbifrost.ai/cli-agents/claude-code) cover the exact variables.

### How do you reduce Claude Code token costs?

Two mechanisms do most of the work at the gateway. [Semantic response caching](https://docs.getbifrost.ai/features/semantic-caching) returns a stored response when a new prompt is semantically equivalent to an earlier one, avoiding a second billed call on short, repeated prompts. [Code Mode for MCP tools](https://docs.getbifrost.ai/mcp/code-mode) has the model write code to orchestrate tools instead of issuing many individual tool calls, which cuts token consumption on MCP-heavy agentic sessions.

## Get Started with Bifrost

Monitoring Claude Code token cost comes down to putting an identity-aware gateway between your developers and your providers, then reading the cost and token metrics it records. Bifrost does this as a free, open-source layer with hierarchical budgets, real-time telemetry, and token-reducing features like semantic caching and Code Mode, and it scales to enterprise governance when you need it. For MCP-heavy Claude Code sessions, the breakdown of [cutting Claude Code token costs with Bifrost as an MCP gateway](https://www.getmaxim.ai/articles/reduce-claude-code-token-costs-by-up-to-90-with-bifrost-mcp-gateway/) shows where the largest savings come from.

Explore the full set of [Bifrost resources](https://www.getmaxim.ai/bifrost/resources) or [book a demo](https://getmaxim.ai/bifrost/book-a-demo) with the Bifrost team to see how LLM gateway cost tracking fits your Claude Code rollout.

## Read next

[![Top 5 AI Gateways for Controlling Shadow AI in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-controlling-shadow-ai-bifrost-isometric.png) Shadow AI is the use of AI tools, models, and MCP servers that security teams have not approved and cannot see. This guide ranks five AI gateways for controlling it, including Bifrost with Bifrost Edge, Kong AI Gateway, Cloudflare AI Gateway, and Gravitee.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-controlling-shadow-ai/)

[![Top 5 AI Gateways for SSO and RBAC in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-sso-and-rbac-in-2026-bifrost-isometric.png) AI gateways with SSO and RBAC let enterprises tie every model request and every configuration change to a corporate identity. This guide compares Bifrost, Kong AI Gateway, Azure API Management, Gravitee, and Cloudflare AI Gateway on identity, roles, provisioning, and audit.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-sso-and-rbac-in-2026/)

[![Semantic Caching: The Top 5 AI Gateways in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/semantic-caching-the-top-5-ai-gateways-in-2026-bifrost-isometric.png) Semantic caching serves a stored LLM response when a new prompt means the same thing as an earlier one. This guide compares Bifrost, Kong AI Gateway, Azure API Management, and Cloudflare AI Gateway on match modes, vector stores, thresholds, TTLs, and cache scoping.](https://www.getmaxim.ai/articles/semantic-caching-the-top-5-ai-gateways-in-2026/)

```json
{
    "@context": "https://schema.org",
    "@type": "Article",
    "publisher": {
        "@type": "Organization",
        "name": "Maxim Articles",
        "url": "https://www.getmaxim.ai/articles/",
        "logo": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w256h256/2025/08/thumbnail.png",
            "width": 60,
            "height": 60
        }
    },
    "author": {
        "@type": "Person",
        "name": "Kuldeep Paul",
        "image": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2025/08/1727978381919.jpeg",
            "width": 800,
            "height": 800
        },
        "url": "https://www.getmaxim.ai/articles/author/kuldeep/",
        "sameAs": []
    },
    "headline": "Best LLM Gateways for Monitoring Claude Code Token Spend",
    "url": "https://www.getmaxim.ai/articles/best-llm-gateways-for-monitoring-claude-code-token-spend/",
    "datePublished": "2026-08-02T18:07:00.000Z",
    "dateModified": "2026-10-08T16:03:59.000Z",
    "image": {
        "@type": "ImageObject",
        "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/best-llm-gateways-for-monitoring-claude-code-token-spend-bifrost-weave.optimized.png",
        "width": 1200,
        "height": 630
    },
    "keywords": "AI Gateway",
    "description": "Claude Code token cost monitoring attributes every request's tokens and dollar cost to a developer, team, or project. This comparison covers Bifrost, LiteLLM, Cloudflare AI Gateway, OpenRouter, and Kong, plus Claude Code's native /usage and OpenTelemetry options.",
    "mainEntityOfPage": "https://www.getmaxim.ai/articles/best-llm-gateways-for-monitoring-claude-code-token-spend/"
}
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "name": "Does routing Claude Code through a gateway change the developer experience?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. Claude Code behaves identically; the only change is the base URL and auth token in settings.json. Model switching with /model and standard workflows continue to work, while the gateway adds cost tracking transparently."
      }
    },
    {
      "@type": "Question",
      "name": "Can a gateway track Claude Code spend per developer?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Yes. By issuing each developer a separate virtual key, Bifrost attributes token usage and cost to the individual key, and rolls those costs up to team and customer levels for reporting."
      }
    },
    {
      "@type": "Question",
      "name": "Which providers can serve Claude models through Bifrost?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Claude models run through Anthropic directly, as well as AWS Bedrock, Google Vertex, and Azure. Bifrost routes Claude Code to any of these, and developers can switch providers mid-session without code changes. The full list is in the supported providers documentation."
      }
    },
    {
      "@type": "Question",
      "name": "Is Bifrost free to use for cost monitoring?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Yes. Virtual keys, budgets, rate limits, telemetry, and cost tracking are part of the open-source build. Enterprise features such as RBAC, SSO, and in-VPC deployment are available when teams need them."
      }
    },
    {
      "@type": "Question",
      "name": "How much does Claude Code cost?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Claude Code billing depends on how it is accessed. Subscription plans bundle usage into a flat monthly fee, while API access bills per input and output token at model-specific rates. Only the API path produces the per-request token costs a gateway can attribute, which is why teams standardizing on API keys route Claude Code through a gateway to see spend per developer rather than one aggregate invoice."
      }
    },
    {
      "@type": "Question",
      "name": "Can Claude Code run through a proxy?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Yes. Claude Code reads its API base URL from an environment variable, so pointing it at a gateway needs no code changes and no change to how developers work. Requests then flow through the gateway, which records tokens and cost per identity before forwarding to Anthropic. The Claude Code integration guide covers the exact variables."
      }
    },
    {
      "@type": "Question",
      "name": "How do you reduce Claude Code token costs?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Two mechanisms do most of the work. Semantic caching returns a stored response when a new prompt is semantically equivalent to an earlier one, avoiding a second billed call. Code Mode has the model write code to orchestrate tools instead of issuing many individual tool calls, which cuts token consumption on MCP-heavy agentic sessions."
      }
    }
  ]
}
```
