Claude Code + Bifrost | Enterprise LLM Gateway for Claude Code

Add multi-provider routing, cost control, guardrails, and governance to Claude Code at scale with Bifrost - the fastest enterprise LLM gateway.

Performance at a Glance

Mean Latency
11µs Gateway overhead per request
Throughput
5K RPS Requests per second sustained
Faster
50x Than Python-based gateways
Providers
25+ Model APIs supported

Pain Points

  • No cost visibility. Teams using Claude Code directly do not get centralized request logs, token attribution, or team-level spend controls.
  • Single provider dependency. A direct Claude Code setup depends on one provider path. Bifrost adds multi-provider routing and fallback policies behind the same developer workflow.
  • No guardrails at scale. Developer prompts and generated code can include sensitive context. Central guardrails help enforce PII redaction and content policies before requests reach models.
  • Limited observability. Without a gateway, platform teams cannot consistently trace Claude Code requests by user, team, route, model, token count, and latency.

Core Features

  • LLM cost control + budgetsCost management and optimization. Track LLM spend per request with breakdowns by provider, model, team, and developer. Virtual keys enforce team-level budgets. Semantic caching reduces costs on repeat queries.
  • 99.999% uptime targetMulti-provider routing with failover. Automatic failover across Anthropic, AWS Bedrock, and Google Vertex AI when rate limits or outages hit. Adaptive load balancing keeps throughput stable even under heavy load.
  • AWS Bedrock + Azure AIGuardrails and governance. Enforce content policies, PII redaction, and safety checks before requests reach the model. Role-based access controls and rate limits per team provide fine-grained LLM governance across the organization.
  • 11µs @ 5K RPSReduce latency, high throughput. Built in Go for production workloads. Bifrost adds only 11µs mean overhead at 5,000 requests per second, making it 50x faster than Python-based gateways. Coding workflows stay fast at scale.
  • OTEL nativeCentralized LLM API observability. Every Claude Code request is logged with full metadata including user, team, provider, route, token count, and latency. Filter and export through the dashboard or push to any observability stack via OpenTelemetry.
  • Vault + SSO readyCentralized credential management. API keys for all providers live in one place. Integrate with HashiCorp Vault for secure key storage or manage them directly in Bifrost. SSO support for Google, GitHub, and enterprise identity providers.

Setup Steps

  1. 01Deploy Bifrost. Bifrost runs as a standalone Go service. Teams deploy it in-VPC or via managed hosting. No agent installation on developer machines.
    Deploy Bifrost
    # pull and start bifrost
    docker pull bifrost-gateway
    docker run -p 8080:8080 bifrost
  2. 02Point Claude Code at it. Developers set one environment variable. Claude Code sends all requests through Bifrost without any code changes or plugin installation.
    Point Claude Code at it
    export ANTHROPIC_BASE_URL="http://localhost:8080"
    # that's it - claude code just works
  3. 03Configure from the dashboard. Set team budgets, apply guardrails, configure provider fallbacks, and view real-time analytics, all from Bifrost's web interface. No code required.
    Configure from the dashboard
    # dashboard available at
    localhost:8080/logs
    # virtual keys, budgets, guardrails

Enterprise Features

  • Automatic failovers. Requests reroute seamlessly when a provider fails or hits rate limits.
  • Adaptive load balancing. Traffic distributes intelligently based on real-time health signals.
  • Semantic caching. Repeat or near-identical queries resolve instantly, cutting costs and reducing latency.
  • Virtual keys & budgets. Create separate virtual API keys for each team with independent limits.
  • Access profiles. Reusable policies for models, budgets, rate limits, and MCP access.
  • User provisioning. SCIM and OIDC from Okta, Microsoft Entra, and other identity providers.
  • Guardrails. Enforce content policies, PII redaction, and safety checks.
  • Audit logs. Complete, tamper-evident record of every request for compliance.
  • Vault support. API keys stored in HashiCorp Vault, never touch developer machines.
  • Cluster mode. Horizontal scaling with zero downtime across multiple nodes.
  • Code Mode (MCP). AI generates Python to orchestrate multiple MCP tools in one execution.
  • Alerts. Threshold-based alerts for cost overruns, rate limits, and errors.
  • MCP Gateway. Inject filesystem tools, database connectors, and custom integrations.
  • Virtual MCPs. Curated tool sets per team, assigned by key, profile, or project.

Comparison Data

FeatureStandaloneWith Bifrost
Multi-model supportNo25+ providers
MCP tool gatewayNoFull MCP injection
Cost trackingNoReal-time per-request
Provider failoverNoAutomatic across providers
Semantic cachingNoReduce costs and latency
Team budgetsNoVirtual keys + limits
Request observabilityNoFull log trail + OTEL export
Gateway latency-11µs at 5,000 RPS

Native MCP Tool Support for Agentic Workflows

Bifrost connects Claude Code to filesystem tools, databases, web search, and custom integrations via Model Context Protocol without modifying the Claude Code client or adding configuration steps on the developer side.

  • Multi-provider development. Teams test code across Claude Sonnet, GPT-4, and Gemini from the same Claude Code workspace. Model performance and cost comparisons happen in real time inside Bifrost's dashboard.
  • Agentic coding pipelines. Claude Code combines with MCP-connected tools for database queries, API testing, deployment scripts, and custom integrations all routed and monitored through a single gateway.
  • Semantic caching at scale. Repeat or near-identical queries across developers resolve instantly from cache. Teams running large codebases see cost savings on common operations like code explanations and documentation generation.

Use Cases

  • Enterprise cost management. Platform teams set department-level budgets for Claude Code usage. Real-time cost tracking surfaces which teams, projects, or developers are driving LLM spend. Automated alerts fire when budgets approach limits.
  • Multi-model testing and comparison. Engineering teams route the same Claude Code workflow through Claude Sonnet, GPT-4, and Gemini to compare code quality, latency, and cost. Bifrost logs performance metrics for each provider.
  • Regulatory compliance and governance. Organizations in healthcare, finance, or government use Bifrost's guardrails to enforce PII redaction and content policies. Audit logs provide tamper-evident records for SOC 2, HIPAA, and GDPR compliance.
  • High-availability production deployments. Teams running Claude Code at scale rely on Bifrost's automatic failover and load balancing to maintain 99.999% uptime. When Anthropic hits rate limits, requests automatically route to Bedrock or Vertex AI.
  • Startups scaling AI development. Early-stage teams use Bifrost's LLM gateway to experiment with multiple providers without vendor lock-in. Semantic caching cuts costs and latency during rapid prototyping.
  • Agentic coding with MCP tools. Developers connect Claude Code to databases, APIs, and deployment pipelines via MCP. Bifrost handles tool injection transparently, enabling automated database migrations and cloud deployment scripts.

Governance Features

  • Role-based access control. Define teams, roles, and environment-specific access at the organization level. Developers, platform engineers, and finance teams each get appropriate visibility and control.
  • Comprehensive audit trails. Every request, policy enforcement action, and configuration change is logged with full context. Export audit trails to your SIEM or compliance platform.
  • Content filtering and PII redaction. Bifrost's guardrails detect and redact sensitive information like SSNs, credit card numbers, and API keys before requests reach the model.
  • In-VPC deployment. Deploy Bifrost entirely within your VPC for maximum security and data control. All LLM requests stay within your network perimeter.

Compliance Badges

  • Audited quarterlySOC 2 Type II. Compliance signal for audit-ready enterprise deployments.
  • EU data residencyGDPR. Data protection signal for privacy and residency reviews.
  • BAA availableHIPAA. Healthcare compliance signal for protected health information workflows.
  • CertifiedISO 27001. Security management signal for enterprise risk reviews.

Open Source & Enterprise

OSS Features

  • 01Model Catalog. Access 8+ providers and 1000+ AI models through a unified interface. Also supports custom deployed models.
  • 02Budgeting. Set spending limits and track costs across teams, projects, and models.
  • 03Provider Fallback. Automatic failover between providers ensures 99.99% uptime for your applications.
  • 04MCP Gateway. Centralize all MCP tool connections, governance, security, and auth. Your AI can safely use MCP tools with centralized policy enforcement. [MCP Gateway resource]
  • 05Virtual Key Management. Create different virtual keys for different use cases with independent budgets and access control.
  • 06Unified Interface. One consistent API for all providers. Switch models without changing code.
  • 07Drop-in Replacement. Replace your existing SDK with just one line change. Compatible with OpenAI, Anthropic, LiteLLM, Google GenAI, LangChain, and more. [Drop-in replacement docs]
  • 08Built-in Observability. Out-of-the-box OpenTelemetry support. Built-in dashboard for quick visibility without complex setup.
  • 09Community Support. Active Discord community with responsive support and regular updates.

Enterprise Features

  • 01Governance. SAML support for SSO and role-based access control with policy enforcement for team collaboration. [Governance resource]
  • 02Adaptive Load Balancing. Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.
  • 03Cluster Mode. High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.
  • 04Alerts. Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook, and more.
  • 05Log Exports. Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export for compliance, monitoring, and analytics.
  • 06Audit Logs. Comprehensive logging and audit trails for compliance and debugging.
  • 07Vault Support. Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.
  • 08VPC Deployment. Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls. [Enterprise deployment resource]
  • 09Guardrails. Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents. [Guardrails resource]

FAQ

How do I configure Claude Code with an LLM gateway?

Claude Code is configured with an [LLM gateway](https://www.getmaxim.ai/llm-gateway) by setting ANTHROPIC_BASE_URL to the gateway's Anthropic-compatible endpoint and ANTHROPIC_AUTH_TOKEN to a gateway credential in settings.json. With Bifrost, the base URL ends in /anthropic and the token is a Bifrost [virtual key](https://www.getmaxim.ai/bifrost/resources/governance), which carries the developer's budget and model access.

What does a Claude Code router do?

A Claude Code router sends Claude Code requests to models other than the default Anthropic endpoint. Bifrost performs this routing at the [AI gateway](https://www.getmaxim.ai/llm-gateway) through provider-prefixed model names and routing rules, and adds budgets, [request logs](https://www.getmaxim.ai/bifrost/ai-observability), and failover to the same path, so routing and [governance](https://www.getmaxim.ai/bifrost/ai-governance) are configured in one service.

How can I monitor my Claude Code usage?

Claude Code usage can be monitored per developer by routing it through Bifrost with one virtual key per person. Every request is logged with tokens, cost, model, and latency, and the same data is available as Prometheus metrics and OpenTelemetry traces for existing dashboards. [AI observability](https://www.getmaxim.ai/bifrost/ai-observability) and [the cost tracking walkthrough](https://docs.getbifrost.ai/features/observability/default) show a full setup.

Do developers still need a Claude Code API key?

Developers using ANTHROPIC_AUTH_TOKEN with a Bifrost virtual key do not need an Anthropic API key or Anthropic account login. Provider credentials stay in Bifrost, and revoking or rotating access is done by changing the virtual key.

Can I use Claude Code with Bedrock or Vertex AI through Bifrost?

Claude Code can use Claude models on [AWS Bedrock](https://www.getmaxim.ai/bifrost/resources/aws-bedrock), Google Vertex AI, and Azure through Bifrost by setting the model to a provider-prefixed value such as bedrock/global.anthropic.claude-sonnet-4-6 or vertex/claude-sonnet-4-6. Billing then runs through your cloud account, and Bifrost can fail over between providers.

How much does Claude Code cost for teams?

Claude Code cost for a team depends on the Anthropic plan or the per-token API pricing of whichever provider serves the requests, and Bifrost does not change those rates. Bifrost records the cost of each request against a developer or team and enforces [budgets](https://www.getmaxim.ai/bifrost/resources/governance), which makes the total predictable and attributable. The open-source gateway is free to run.