Use Bifrost to scale Claude Code across your organization with multi-provider routing on one AI gateway, cost controls, security guardrails, role-based access control, and compliance-ready governance.
[ PERFORMANCE AT A GLANCE ]
[ THE PROBLEM ]
Claude Code is powerful out of the box for individual developers. But scaling it across an engineering organization surfaces problems that Anthropic doesn't solve.
No way to track which teams, projects, or developers are driving Claude Code spend. Budgets are managed manually.
When Anthropic hits rate limits or has an outage, every developer using Claude Code stops working. No fallback, no failover.
Sensitive data, PII, and internal code flow freely through the API. No content policies, no redaction, no audit trail for compliance.
No centralized view of requests, token usage, latency, or error rates. Platform teams fly blind when rolling out AI tooling org-wide.
[ HOW IT WORKS ]
Set one environment variable to route Claude Code through Bifrost, developers work unchanged while platform teams gain full control over budgets, guardrails, failover routing, and real-time observability across 25+ providers.
Nothing changes. Set one environment variable and Claude Code works exactly as before. Same API, same workflow, same speed.
Full control. Set budgets per team, enforce guardrails, configure failover routes, and get real-time observability across every Claude Code request in the organization.
[ COST AND USAGE ]
Bifrost attributes Claude Code cost and token usage to the virtual key that sent each request, and enforces limits at the same point.
Budgets and rate limits apply independently at the customer, team, virtual key, and provider-config level. A request proceeds only if every applicable budget has balance, with resets daily, weekly, monthly, quarterly, or yearly.
Request logs capture input, output, tokens, cost, and latency for every request, written asynchronously so logging adds no request latency. Filter by token and latency range in the dashboard or API.
Bifrost exposes Prometheus metrics at /metrics and exports traces through OpenTelemetry, so Claude Code monitoring lands in the dashboards a platform team already runs.
Bifrost Enterprise alert rules notify Slack, Microsoft Teams, PagerDuty, or a webhook when a team or virtual key crosses a budget or rate-limit threshold.
[ SETUP ]
No SDK changes, no plugin installation, no developer workflow disruption.
Bifrost runs as a standalone Go service. Teams deploy it in-VPC or via managed hosting. No agent installation on developer machines.
Developers set one environment variable. Claude Code sends all requests through Bifrost without any code changes or plugin installation.
Set team budgets, apply guardrails, configure provider fallbacks, and view real-time analytics, all from Bifrost's web interface. No code required.
[ COMPARISON ]
| Feature | Claude Code (standalone) | Claude Code + Bifrost |
|---|---|---|
| Multi-model support | No | 25+ providers |
| MCP tool gateway | No | Full MCP injection |
| Cost tracking | No | Real-time per-request |
| Provider failover | No | Automatic across providers |
| Semantic caching | No | Reduce costs and latency |
| Team budgets | No | Virtual keys + limits |
| Request observability | No | Full log trail + OTEL export |
| Gateway latency | N/A | 11µs at 5,000 RPS |
[ MODELS AND PROVIDERS ]
Claude Code can call any model Bifrost is configured for by prefixing the model name with its provider, as long as that model supports the tool calling Claude Code relies on for file edits and shell commands.
| Target | Example model value | Notes |
|---|---|---|
| Anthropic | claude-sonnet-4-6 | No prefix required |
| AWS Bedrock | bedrock/global.anthropic.claude-sonnet-4-6 | Claude on your AWS account |
| Google Vertex AI | vertex/claude-sonnet-4-6 | Claude on your GCP project |
| Azure | azure/claude-sonnet-4-6 | Verify tool-use support first |
| OpenAI | openai/gpt-5.5 | Claude-specific server tools unavailable |
| Gemini on Vertex AI | vertex/gemini-3.1-pro | Claude-specific server tools unavailable |
The /model command and the --model flag accept any provider-prefixed model, and Claude Code continues the conversation context on the new model.
Routing rules can map an alias such as sonnet-model to a different provider per team or header, so a model change is a gateway edit and not a change on every developer machine.
Retries and fallbacks retry transient errors within a provider and switch to a configured backup provider when retries are exhausted, which keeps Claude Code sessions running through provider rate limits.
[ MCP TOOLS ]
Bifrost aggregates connected MCP servers behind a single /mcp endpoint, so Claude Code needs one MCP entry for the whole tool set, with per-key tool filtering deciding which tools each developer sees.
claude mcp add --transport http bifrost http://localhost:8080/mcp --header "Authorization: Bearer your-virtual-key" --scope user
For large tool sets, Code Mode exposes four generic tools in place of the full catalog and reduces input tokens by up to 92.8% across multiple MCP servers. The MCP gateway page covers authentication modes and tool governance.
[ ENTERPRISE CONTROLS ]
Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, and a Claude Code enterprise rollout uses the same controls as any other workload on the AI gateway.
Users sign in through OIDC providers such as Okta and Microsoft Entra, with SCIM provisioning for teams and roles. Access profiles issue each user a virtual key automatically, carrying the profile's models, budgets, rate limits, and MCP access.
Guardrails inspect inputs before the model call and outputs after it. Secrets detection flags API keys, tokens, and private keys in prompts, which matters when an agent reads configuration files.
Audit logs record administrative activity such as key and policy changes, alongside the request logs above. Bifrost runs in your VPC or air-gapped, with clustering for high availability. Bifrost Enterprise includes a 14-day trial.
The gateway is the policy engine, and Bifrost Edge, currently in alpha, extends the same virtual keys, budgets, and guardrails to Claude Code on every company machine without per-developer configuration.
[ BUILT FOR PRODUCTION ]
Bifrost ships with the full set of controls platform teams expect before rolling out AI tooling organization-wide.
Requests reroute seamlessly when a provider fails or hits rate limits.
Traffic distributes intelligently based on real-time health signals.
Repeat or near-identical queries resolve instantly, cutting costs and reducing latency.
Create separate virtual API keys for each team with independent limits.
Reusable policies for models, budgets, rate limits, and MCP access.
SCIM and OIDC from Okta, Microsoft Entra, and other identity providers.
Enforce content policies, PII redaction, and safety checks.
Complete, tamper-evident record of every request for compliance.
API keys stored in HashiCorp Vault, never touch developer machines.
Horizontal scaling with zero downtime across multiple nodes.
AI generates Python to orchestrate multiple MCP tools in one execution.
Threshold-based alerts for cost overruns, rate limits, and errors.
Inject filesystem tools, database connectors, and custom integrations.
Curated tool sets per team, assigned by key, profile, or project.
[ AGENTIC WORKFLOWS ]
Bifrost connects Claude Code to filesystem tools, databases, web search, and custom integrations via Model Context Protocol without modifying the Claude Code client or adding configuration steps on the developer side.
Teams test code across Claude Sonnet, GPT-4, and Gemini from the same Claude Code workspace. Model performance and cost comparisons happen in real time inside Bifrost's dashboard.
Claude Code combines with MCP-connected tools for database queries, API testing, deployment scripts, and custom integrations all routed and monitored through a single gateway.
Repeat or near-identical queries across developers resolve instantly from cache. Teams running large codebases see cost savings on common operations like code explanations and documentation generation.
[ USE CASES ]
Platform teams set department-level budgets for Claude Code usage. Real-time cost tracking surfaces which teams, projects, or developers are driving LLM spend. Automated alerts fire when budgets approach limits.
Engineering teams route the same Claude Code workflow through Claude Sonnet, GPT-4, and Gemini to compare code quality, latency, and cost. Bifrost logs performance metrics for each provider.
Organizations in healthcare, finance, or government use Bifrost's guardrails to enforce PII redaction and content policies. Audit logs provide tamper-evident records for SOC 2, HIPAA, and GDPR compliance.
Teams running Claude Code at scale rely on Bifrost's automatic failover and load balancing to maintain 99.999% uptime. When Anthropic hits rate limits, requests automatically route to Bedrock or Vertex AI.
Early-stage teams use Bifrost's LLM gateway to experiment with multiple providers without vendor lock-in. Semantic caching cuts costs and latency during rapid prototyping.
Developers connect Claude Code to databases, APIs, and deployment pipelines via MCP. Bifrost handles tool injection transparently, enabling automated database migrations and cloud deployment scripts.
[ GOVERNANCE & COMPLIANCE ]
Bifrost ships with the governance features and compliance certifications platform teams need before rolling out AI tooling organization-wide.
Define teams, roles, and environment-specific access at the organization level. Developers, platform engineers, and finance teams each get appropriate visibility and control.
Every request, policy enforcement action, and configuration change is logged with full context. Export audit trails to your SIEM or compliance platform.
Bifrost's guardrails detect and redact sensitive information like SSNs, credit card numbers, and API keys before requests reach the model.
Deploy Bifrost entirely within your VPC for maximum security and data control. All LLM requests stay within your network perimeter.
[ WHAT'S NEXT ]
Continue with governance, guardrails, MCP, and the rest of the resource library.
[ BIFROST FEATURES ]
Everything you need to run AI in production, from free open source to enterprise-grade features.
01 Governance
SAML support for SSO and Role-based access control and policy enforcement for team collaboration.
02 Adaptive Load Balancing
Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.
03 Cluster Mode
High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.
04 Alerts
Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook and more.
05 Log Exports
Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export capabilities for compliance, monitoring, and analytics.
06 Audit Logs
Comprehensive logging and audit trails for compliance and debugging.
07 Vault Support
Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.
08 VPC Deployment
Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls.
09 Guardrails
Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents.
[ SHIP RELIABLE AI ]
Change just one line of code. Works with OpenAI, Anthropic, Vercel AI SDK, LangChain, and more.
[ FAQ ]
Claude Code is configured with an LLM gateway by setting ANTHROPIC_BASE_URL to the gateway's Anthropic-compatible endpoint and ANTHROPIC_AUTH_TOKEN to a gateway credential in settings.json. With Bifrost, the base URL ends in /anthropic and the token is a Bifrost virtual key, which carries the developer's budget and model access.
A Claude Code router sends Claude Code requests to models other than the default Anthropic endpoint. Bifrost performs this routing at the AI gateway through provider-prefixed model names and routing rules, and adds budgets, request logs, and failover to the same path, so routing and governance are configured in one service.
Claude Code usage can be monitored per developer by routing it through Bifrost with one virtual key per person. Every request is logged with tokens, cost, model, and latency, and the same data is available as Prometheus metrics and OpenTelemetry traces for existing dashboards. AI observability and the cost tracking walkthrough show a full setup.
Developers using ANTHROPIC_AUTH_TOKEN with a Bifrost virtual key do not need an Anthropic API key or Anthropic account login. Provider credentials stay in Bifrost, and revoking or rotating access is done by changing the virtual key.
Claude Code can use Claude models on AWS Bedrock, Google Vertex AI, and Azure through Bifrost by setting the model to a provider-prefixed value such as bedrock/global.anthropic.claude-sonnet-4-6 or vertex/claude-sonnet-4-6. Billing then runs through your cloud account, and Bifrost can fail over between providers.
Claude Code cost for a team depends on the Anthropic plan or the per-token API pricing of whichever provider serves the requests, and Bifrost does not change those rates. Bifrost records the cost of each request against a developer or team and enforces budgets, which makes the total predictable and attributable. The open-source gateway is free to run.