Try Bifrost Enterprise free for 14 days. Request access

Top 5 AI Gateways for Claude Code, Codex CLI, and Cursor

AI gateways sit between coding agents and model providers so platform teams can issue per-developer keys, cap spend, and switch models centrally. This guide compares Bifrost, LiteLLM, Vercel AI Gateway, OpenRouter, and Kong AI Gateway for Claude Code, Codex CLI, and Cursor traffic.

Top 5 AI Gateways for Claude Code, Codex CLI, and Cursor

TL;DR

  • AI gateways for coding agents must expose an Anthropic-format endpoint for Claude Code, an OpenAI Responses-compatible endpoint for Codex CLI, and an OpenAI-compatible base URL for Cursor.
  • Bifrost documents integrations for Claude Code, Codex CLI, Cursor, Gemini CLI, Opencode, and more, with per-developer virtual keys, budgets, and MCP tool filtering on one gateway.
  • Running Claude Code against non-Anthropic models works through gateway model mapping, but the target model must support the tool calls Claude Code relies on.
  • Bifrost adds 11 microseconds of overhead per request at 5,000 RPS and reaches 10,000+ models across 25+ providers through one API.
  • AI Gateway + Bifrost Edge (alpha) extends the same gateway policies to coding agents on developer laptops that were never configured to use the gateway.

Stack Overflow's Developer Survey found that 84% of respondents use or plan to use AI tools in their development process, and in engineering organizations a growing share of that usage runs through Claude Code, Codex CLI, and Cursor. AI gateways give platform teams one place to issue per-developer credentials, cap spend, switch models, and govern the MCP tools those agents call. Bifrost, the open-source AI gateway written in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, and it ships documented integrations for all three agents. This guide compares five AI gateways on how well they handle coding-agent traffic specifically.

What an AI Gateway Does for Coding Agents

An AI gateway for coding agents is a control layer that receives every model and tool request from Claude Code, Codex CLI, or Cursor, authenticates the developer, enforces budgets, routes to a provider, and logs the result. Developers keep their tools; the platform team gains one policy point for every session.

The hard part is protocol coverage. Each agent speaks a different API shape, so a gateway that only offers an OpenAI-compatible endpoint covers Cursor but not Claude Code. Figure 1 shows the three inbound formats that have to converge on one policy layer.

Claude Code, Codex CLI, and Cursor each send traffic in their native API format to one AI gateway, which applies per-developer keys and routes to model providers and MCP servers

Figure 1: The gateway has to speak each agent's native API format before it can apply one policy to all three.

Anthropic's own Claude Code gateway documentation lists what a gateway provides: server-side provider credentials, usage attribution by developer or team, budgets and rate limits in one place, audit logging, and provider switching without touching developer machines. For a deeper walkthrough of the Claude-specific pieces, see how a Claude Code gateway handles routing, governance, and cost control.

Key Criteria for Choosing AI Gateways for Coding Agents

The criteria that matter for coding-agent traffic differ from those for application traffic. Coding agents run long, tool-heavy sessions per developer, so per-user identity, spend caps, prompt-cache stickiness, and MCP tool control matter more than raw request throughput or semantic caching.

Criterion Why it matters for coding agents What to check
Native agent endpoints Claude Code, Codex CLI, and Cursor each expect a different API format Anthropic Messages, OpenAI Responses, and OpenAI-compatible endpoints on one gateway
Per-developer keys Usage must be attributable and revocable per engineer Virtual or scoped keys that replace raw provider keys on laptops
Budgets and rate limits One long agent session can consume a large token volume Dollar budgets with reset periods plus token and request limits per key
Model switching Teams route Claude Code to Bedrock, Vertex AI, or other model families Model aliasing or mapping without editing every developer's config
Usage visibility Finance and platform teams need cost per developer and per model Request logs with tokens, cost, latency, and filters
MCP tool governance Agents call tools that read code, open tickets, and hit internal APIs Per-key tool allow lists and a single MCP endpoint
Deployment control Source code passes through the gateway Self-hosted, in-VPC, or on-prem options

The Bifrost governance resource page covers how these controls map to virtual keys, teams, and customers inside one AI gateway.

AI Gateways for Claude Code, Codex CLI, and Cursor Compared

The five AI gateways below all accept coding-agent traffic, but they differ in which agents are documented, how per-developer budgets work, and whether MCP tools are governed at the gateway. Cells marked "Not published" mean the vendor page read for this comparison did not state the capability.

Gateway Claude Code Codex CLI Cursor Per-developer keys and budgets Non-Anthropic models in Claude Code MCP tool governance Deployment
Bifrost Documented (/anthropic) Documented (/openai/v1, Responses) Documented (base URL override) Virtual keys with budgets, rate limits, model filters Documented via aliases and provider pinning Per-virtual-key tool allow lists, /mcp endpoint, Code Mode Self-hosted, in-VPC, on-prem, air-gapped
LiteLLM Documented Documented (custom provider) Best-effort guidance Virtual keys with max_budget and TPM/RPM limits Documented via model mapping MCP gateway with key, team, and org permissions Self-hosted proxy
Vercel AI Gateway Dedicated endpoint Dedicated endpoint Dedicated endpoint Budgets per team, project, API key, and user Model picker via gateway discovery Not published Hosted
OpenRouter Documented Documented Not published Team budget management and usage by member Only Anthropic first-party guaranteed Not published Hosted
Kong AI Gateway Documented how-to guides Named as supported Not published Token quotas and cost budgets per consumer group Documented via AI Proxy format translation MCP tool exposure and OAuth2 scoping Konnect or self-hosted Enterprise

Teams that already standardized on one agent can go deeper with the comparison of AI gateways for governing Claude Code and Codex CLI.

1. Bifrost

Bifrost is an open-source AI gateway that exposes Anthropic-, OpenAI-, and Gemini-compatible endpoints, so Claude Code, Codex CLI, Cursor, Gemini CLI, and other agents connect by changing a base URL and supplying a virtual key. Governance, model routing, MCP tool access, and logging then apply to every agent session.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

The Bifrost gateway documents a setup page for each agent in its CLI agents and editors guide, covering Claude Code, Codex CLI, Cursor, Gemini CLI, Qwen Code, Opencode, Zed, Roo Code, GitHub Copilot, and Claude Desktop. Figure 2 shows what happens to one developer's request.

A coding agent request carrying a developer virtual key passes through Bifrost for key checks, budget and rate limits, and routing rules before reaching a provider and being logged

Figure 2: Every check runs against the developer's virtual key, so one key decides models, spend, and tools.

Agent setup for Claude Code, Codex CLI, and Cursor

  • Claude Code: set ANTHROPIC_BASE_URL to the gateway's /anthropic path and ANTHROPIC_AUTH_TOKEN to a virtual key. With that method, no Anthropic account login is required. The Claude Code integration also covers /model switching mid-session.
  • Codex CLI: add a named model_providers entry pointing at /openai/v1 with wire_api = "responses" and export the virtual key as OPENAI_API_KEY. The Codex CLI setup explains why non-OpenAI models need supports_websockets = false and how to add them to the /models picker.
  • Cursor: enter a virtual key in the OpenAI API Key field and enable Override OpenAI Base URL. The Cursor configuration notes that Cursor needs a publicly reachable Bifrost URL and lets teams assign different models to Chat, Agent, and Inline Edit.

For terminal agents, the Bifrost CLI (npx -y @maximhq/bifrost-cli) launches Claude Code, Codex CLI, Gemini CLI, or Opencode with base URLs, keys, and models set automatically, and stores the virtual key in the OS keyring rather than on disk.

Per-developer virtual keys and budgets

Virtual keys are the primary governance entity in Bifrost. Each key carries allowed providers and models, dollar budgets with daily, weekly, monthly, quarterly, or yearly resets (optionally calendar-aligned), and token and request rate limits. Budgets and limits stack across virtual key, team, and customer levels, and a request proceeds only if every applicable budget has room.

At enterprise scale, access profiles auto-issue a write-protected virtual key to each user based on role, with isolated budget and rate-limit counters per person.

Claude Code usage and cost visibility

Every agent request is logged with inputs, outputs, tokens, cost, and latency in built-in observability, filterable by provider, model, status, cost range, and content. Claude Code and Codex CLI both send a session header that Bifrost uses for session affinity, so a session keeps hitting the same provider key and prompt cache.

MCP tool governance for coding agents

Bifrost acts as an MCP gateway that aggregates connected MCP servers behind one /mcp endpoint, which Claude Code adds with a single claude mcp add command. MCP tool filtering is deny-by-default per virtual key.

Code Mode replaces large tool catalogs with four meta-tools, cutting input tokens by up to 92.8% in benchmarks. The MCP gateway benchmark writeup shows how those savings grow with tool count.

Bifrost adds 11 microseconds of overhead per request at 5,000 RPS in sustained benchmarks, which keeps the gateway out of the latency budget of interactive agent sessions. One virtual key can reach 10,000+ models across 25+ providers through the same OpenAI-compatible API.

2. LiteLLM

LiteLLM is a self-hosted proxy that documents Claude Code, Codex CLI, and Cursor setups, using virtual keys to restrict models and track spend per key, user, and team. It fits teams already running the LiteLLM proxy who want to add coding-agent traffic to an existing deployment.

Best for: Teams with an existing LiteLLM proxy that want Claude Code and Codex CLI on the same keys and spend tracking.

For Claude Code, LiteLLM's tutorial points ANTHROPIC_BASE_URL at the proxy and uses ANTHROPIC_AUTH_TOKEN with a master key or a virtual key, where a virtual key is limited to the models it has access to. Model mapping in config.yaml sends Claude Code requests to Bedrock, Azure, or Vertex AI deployments.

  • Keys and budgets: virtual keys support max_budget, budget_duration, and TPM and RPM limits, and spend is recorded against a key, user, or team.
  • Codex CLI: configured as a custom model provider in ~/.codex/config.toml, where the MCP gateway is registered too.
  • Cursor: LiteLLM describes its Cursor guidance as best-effort, since Cursor does not officially support AI gateways, and states that the Cursor CLI cannot connect.
  • MCP: the proxy exposes MCP servers through one endpoint with permissions scoped by key, team, or organization.

Teams comparing the two gateways on performance and governance can review Bifrost as a LiteLLM alternative.

3. Vercel AI Gateway

Vercel AI Gateway is a hosted gateway with dedicated compatibility endpoints for Claude Code, Codex CLI, and Cursor, plus a CLI command that writes agent configuration automatically. It suits teams already on Vercel that want hosted billing and spend caps without operating gateway infrastructure.

Best for: Teams on Vercel that want a hosted gateway with one-command agent setup and per-user spend caps.

Vercel documents roughly 30 coding agents, with vercel ai-gateway setup detecting installed agents and writing their config. Claude Code, Codex, and Cursor each get a dedicated endpoint that adapts the agent's request format.

  • Budgets: caps apply at team, project, API key, and user scope, with daily, weekly, monthly, or no reset. Vercel describes a budget as a soft cap, and BYOK provider spend is not counted toward budgets.
  • Cursor limitations: Vercel notes Cursor-side constraints, including that Tab completion never uses a custom key and that bring-your-own-key traffic passes through Cursor's backend.
  • Deployment: hosted only, so prompts and source code transit Vercel infrastructure.

Organizations that need the same coding-agent controls inside their own network can compare self-hosted options in the guide to a self-hosted AI gateway for Cursor.

4. OpenRouter

OpenRouter is a hosted model router with setup guides for Claude Code and Codex CLI, team budget management, and an activity dashboard that breaks usage down by project and team member. Its Claude Code guidance is the most conservative of the five on model switching.

Best for: Individual developers and small teams that want hosted access to many models with credit-based billing.

For Claude Code, OpenRouter sets ANTHROPIC_BASE_URL to its API and the OpenRouter key as ANTHROPIC_AUTH_TOKEN. Its documentation states that Claude Code through OpenRouter is only guaranteed to work with the Anthropic first-party provider, and it uses routing across multiple Anthropic providers for failover.

  • Codex CLI: configured in config.toml with OpenRouter model slugs and command-based auth so Codex can fetch the model catalog.
  • Spend: team-level spending limits and credit allocation, with usage visible in the Activity Dashboard.
  • Cursor and MCP: neither was covered in the OpenRouter pages read for this comparison.

In Bifrost compatibility testing for Claude Code, OpenRouter used as an upstream provider did not stream function-call arguments correctly at the time of writing, which breaks Claude Code file operations. The Bifrost resource page for Claude Code covers the self-hosted alternative.

5. Kong AI Gateway

Kong AI Gateway extends Kong's API gateway with AI Proxy and AI policies for LLM, MCP, and agent-to-agent traffic. Kong publishes how-to guides that route Claude Code through its AI Proxy to OpenAI, Bedrock, Vertex AI, Azure, and Gemini, and names Codex CLI as a supported command-line tool.

Best for: Organizations already running Kong Gateway that want coding-agent traffic under existing API management.

Kong's Claude Code guide sets llm_format: anthropic on the AI Proxy so the gateway accepts Claude's native format and forwards to a non-Anthropic model, with a File Log plugin capturing model, latency, and token usage. That tutorial requires Kong Gateway Enterprise or a Konnect control plane.

  • Spend controls: token quotas and cost budgets are set per AI consumer group and enforced by a rate-limiting policy.
  • MCP: existing APIs can be exposed as MCP tools, with aggregation and OAuth2 scoping for tool calls.
  • Cursor: not mentioned on the Kong AI Gateway pages read for this comparison.

Teams weighing plugin-based API gateways against purpose-built AI gateways can use the LLM gateway buyer's guide as a checklist.

Running Claude Code With Bedrock, Vertex AI, and Non-Anthropic Models

AI gateways let Claude Code reach models beyond Anthropic's API by keeping the Anthropic request format at the client and changing the provider and model at the gateway. Teams use this to run Claude on Bedrock or Vertex AI under existing cloud contracts, or to test other model families on specific tasks.

In Bifrost, two methods are documented. Provider pinning sets values such as bedrock/global.anthropic.claude-sonnet-4-6 or openai/gpt-5.5 directly in Claude Code's settings.json. Dynamic aliasing sends a label such as sonnet-model, and a routing rule matched on the model name and the claude-cli user agent rewrites it at request time, as Figure 3 shows.

Claude Code sends a model alias to the Bifrost Anthropic endpoint, a routing rule rewrites the alias, and the request reaches Bedrock, Vertex AI, or an OpenAI model

Figure 3: Claude Code keeps speaking the Anthropic format while the gateway decides which provider and model serve it.

Two caveats apply.

Anthropic states that it does not support routing Claude Code to non-Claude models through any gateway, and Claude-specific server-side tools such as web search and computer use are available only on Claude-family models. Any pinned model must support the tool calls Claude Code uses for file edits and bash. The model-mapping patterns are compared across gateways in this guide to routing, governance, and cost control for Claude Code. The same provider/model pattern works in Codex CLI, as covered in running Codex CLI with Claude, Gemini, and other models.

Coding Agents That Never Reach the Gateway: AI Gateway + Bifrost Edge

AI gateways only govern traffic that is configured to flow through them. A developer who installs Claude Code with a personal key, or wires a new MCP server into Cursor, bypasses every budget and allow list. AI Gateway + Bifrost Edge closes that gap: the gateway stays the policy engine, and Edge extends it to the laptop.

The Bifrost AI gateway remains where virtual keys, budgets, guardrails, and logs are configured. Bifrost Edge, currently in alpha, runs on each machine and routes AI traffic through that gateway with no per-app base URL changes, so the same policies apply to sessions nobody configured.

Claude Code, Codex CLI, and Cursor on a developer laptop route through Bifrost Edge to the Bifrost AI gateway, which holds virtual keys, budgets, guardrails, and logs

Figure 4: The gateway stays the policy engine; Edge brings agents that were never configured under that same policy.

Edge's supported applications include Claude Code, Codex CLI, and OpenCode as coding agents, and Cursor and the Codex desktop app as desktop apps. Edge MCP governance inventories the MCP servers configured inside Claude Code, Codex, Cursor, and Gemini CLI across the fleet, and a denied server is blocked on the device.

Teams can roll Edge out through MDM platforms such as Jamf, Intune, and Kandji. The broader case is laid out in extending AI governance from the gateway to the endpoint, and the agent-specific policy design in governing Claude Code, Cursor, and Codex at scale.

Frequently Asked Questions

What does a Claude Code router do?

A Claude Code router sits between Claude Code and model providers and decides which provider and model serve each request. In an AI gateway such as Bifrost, routing rules can rewrite a model alias sent by Claude Code to Claude on Bedrock, Vertex AI, or another tool-calling model, while the gateway also enforces the developer's virtual key, budget, and rate limits.

What LLM does Claude Code use?

Claude Code uses Anthropic's Claude models by default, with Sonnet, Opus, and Haiku tiers for different task types. When Claude Code points at a gateway through ANTHROPIC_BASE_URL, the gateway decides which provider serves each tier. Bifrost lets teams pin each tier to Claude on Anthropic, Bedrock, Vertex AI, or Azure, or to another provider's model that supports tool calling.

What does Claude Code connect to?

Claude Code connects to a model API for inference and to MCP servers for tools. By default the model API is Anthropic's; with a gateway, ANTHROPIC_BASE_URL redirects inference to the gateway. Bifrost can also serve as the single MCP server Claude Code connects to, exposing only the tools each developer's virtual key allows through one /mcp endpoint.

How do teams track Claude Code usage per developer?

Teams track Claude Code usage per developer by issuing each engineer a separate gateway credential instead of a shared provider key. In Bifrost, each developer gets a virtual key, and every request is logged with tokens, cost, latency, provider, and model. Budgets and rate limits on the same key cap spend before an overrun, and team-level budgets roll up across developers.

Can Codex CLI use Claude or Gemini models through an AI gateway?

Codex CLI can use Claude, Gemini, and other models through an AI gateway that translates the OpenAI Responses API to other providers. With Bifrost, developers start Codex with --model anthropic/... or --model gemini/..., and non-OpenAI models require HTTPS mode because Codex's WebSocket mode expects server-held conversation state. The model must support tool use for file and terminal operations.

Does Cursor work with an AI gateway?

Cursor works with an AI gateway through its OpenAI API key field and the Override OpenAI Base URL setting. With Bifrost, developers paste a virtual key, set the gateway URL, and add models in provider/model format for Chat, Agent, and Inline Edit. Cursor requires a publicly reachable gateway URL, and non-native models must support tool use for agent mode.

Try Bifrost Today

AI gateways turn Claude Code, Codex CLI, and Cursor from individually configured tools into governed, observable traffic with per-developer keys, budgets, model routing, and MCP tool control. Bifrost covers all three agents with documented setups, adds 11 microseconds of overhead per request, and runs self-hosted, in-VPC, or on-prem through Bifrost Enterprise. Explore the coding agent integrations, or book a demo with the Bifrost team to plan a coding-agent gateway rollout.