Control who can access which models, what budget limits apply, and which MCP tools they can use from a single policy layer on the AI gateway.
[ HOW GOVERNANCE WORKS ]
Each request is authenticated with a virtual key, checked against that key's access rules and every budget above it, and routed only to providers that still have capacity.
The caller sends an sk-bf-* key in x-bf-vk, Authorization: Bearer, x-api-key, x-goog-api-key, or api-key. A request with no key is rejected with a 401 when virtual keys are mandatory, and an inactive or expired key returns a 403.
Bifrost validates the requested provider and model against the key's provider configurations. A key with no provider configurations blocks all providers, and an empty allowed_models list blocks all models for that provider.
Every applicable budget must have remaining balance, and the request must pass both the token and request limits at the provider config and virtual key levels. A spent budget returns a 402 and an exhausted rate limit returns a 429.
Traffic is split across allowed providers by weight, skipping any provider over its own limits, and the request's cost is deducted from every budget in the hierarchy.
The configuration below creates a team-scoped virtual key with two providers, a monthly budget, and token and request limits:
curl -X POST http://localhost:8080/api/governance/virtual-keys \
-H "Content-Type: application/json" \
-d '{
"name": "Engineering Team API",
"provider_configs": [
{ "provider": "openai", "weight": 0.5, "allowed_models": ["gpt-4o-mini"], "key_ids": ["*"] },
{ "provider": "anthropic", "weight": 0.5, "allowed_models": ["claude-3-sonnet-20240229"], "key_ids": ["*"] }
],
"team_id": "team-eng-001",
"budgets": [{ "max_limit": 100.00, "reset_duration": "1M" }],
"rate_limit": {
"token_max_limit": 10000, "token_reset_duration": "1h",
"request_max_limit": 100, "request_reset_duration": "1m"
},
"is_active": true
}'[ OSS GOVERNANCE ]
These controls ship in the open-source gateway and are configured through the Web UI, the REST API under /api/governance/*, config.json, or the Helm chart.
Each key carries model access, budgets, rate limits, and status, and belongs to a user, team, customer, or business unit. Optional expiry rejects the key after a set time without deleting it.
Independent limits at customer, team, virtual key, and provider levels, with the same cost deducted from each. Reset from one day to one year, with optional UTC calendar alignment.
Cap tokens and requests at the virtual key and provider levels, with 1-minute, 1-hour, or 1-day windows. A provider over its limit is skipped; others on the same key stay available.
Split traffic across providers by weight and bind each to specific API key IDs. Separate credentials and spend for development, staging, and production.
Reject LLM or MCP requests missing a configured header, such as a tenant or correlation ID, with a 400 before they reach a provider.
[ VIRTUAL KEYS ]
Virtual keys authenticate requests and enforce access control, budgets, and rate limits per consumer, including securing AI agents with virtual keys and tool filtering.
Restrict which AI models users can access
Limit access to specific AI providers
Independent cost tracking per virtual key
Token and request-based throttling
Bind to specific provider keys
Instantly enable or disable access
x-bf-vk
sk-bf-*
Native Bifrost
Authorization
Bearer
OpenAI-compatible
x-api-key
sk-ant-*
Anthropic-compatible
x-goog-api-key
AI*
Gemini-compatible
[ INTELLIGENT ROUTING ]
Distribute traffic across providers with configurable weights and automatic fallback chains.
Automatically optimizes traffic distribution across providers and keys based on real-time performance metrics.
Create fallback chains ordered by weight when primary providers fail or hit rate limits
Whitelist specific provider-model combinations with empty array defaulting to catalog detection
Restrict VKs to specific provider API keys for environment separation (dev/test/prod)
[ HIERARCHICAL BUDGETS ]
Independent cost tracking at Customer, Team, Virtual Key, and Business Unit levels with automatic deduction across all tiers.
Top-level organization with independent budget
Department-level budget within customer
Individual access token budget
Business-unit budget across teams
All Applicable Budgets Must Pass
When a transaction occurs, the same cost deducts from every relevant level simultaneously. A single exhausted budget at any tier blocks the entire request.
[ MCP TOOL FILTERING ]
AI agent governance in Bifrost uses one virtual key to define an agent's model access, spend, and MCP tool permissions, so agent traffic counts against the same budgets and rate limits as the team that owns it. Tool connectivity itself is covered on the MCP gateway page.
Tools stay unavailable to a virtual key until an MCP client is configured on it, except clients marked Allow by Default. Each client can allow selected tools, all tools via *, or none.
Bifrost generates an x-bf-mcp-include-tools header from the key's allow-list. Caller-supplied entries the key does not allow are dropped, and the allow-list is enforced again when the tool executes.
[ ENTERPRISE GOVERNANCE ]
Bifrost Enterprise adds identity and administration controls to the open-source policy engine.
| Capability | Open source | Enterprise |
|---|---|---|
| Virtual keys, budgets, rate limits, routing | ||
| MCP tool filtering and required headers | ||
| RBAC with system and custom roles | ||
| SSO through OIDC and SCIM 2.0 provisioning | ||
| Access Profiles and per-user virtual keys | ||
| Signed audit logs and budget alerting |
RBAC ships with Admin, Developer, and Viewer system roles plus custom roles, assigned automatically from identity provider groups and claims.
Connect Okta, Microsoft Entra, Keycloak, Zitadel, Google Workspace, Auth0, or any standards-compliant OIDC provider. Teams and business units sync from the IdP; when a user matches several roles, the highest-privilege role applies.
Reusable policies covering providers, models, budgets, rate limits, and MCP access. Assigning a profile issues each user a write-protected virtual key with isolated counters; profile edits apply on the next request.
HMAC-signed administrative events that can be archived to S3-compatible storage. Budget and rate-limit alerts go to Slack, Microsoft Teams, PagerDuty, or webhooks.
Deployment options are listed on the Bifrost Enterprise page.
[ COMPLIANCE ]




[ ROLE-BASED ACCESS CONTROL ]
Start with Admin, Developer, and Viewer. Create custom roles as your team structure evolves.
Users receive only the permissions they need for their job function, reducing security vulnerabilities and preventing accidental misconfigurations.
Assign roles once instead of configuring individual permissions. New team members inherit appropriate access automatically through role assignment.
Demonstrate to auditors exactly who has what access. Audit logs track permission changes over time for compliance frameworks like SOC 2 Type II and GDPR.
Create tailored roles for QA teams, security auditors, or compliance officers. Custom roles adapt to your organizational structure.
Full control over all Bifrost resources and configurations
Platform engineers, security adminsManage technical resources without administrative privileges
Engineering teams, DevOpsRead-only access for monitoring and compliance
Finance, compliance, executives[ CONFIGURATION ]
Visual dashboard for configuring virtual keys, budgets, routing, and RBAC
Programmatic management via endpoints at /api/governance/*
Declarative file-based configuration for GitOps workflows
Interactive terminal setup for managing governance settings from your workflow
[ USE CASES ]
Isolate tenants with virtual keys, enforce per-tenant budgets, and track usage with required headers. Automatic cost allocation across customers.
Department-level budgets with team-specific provider access. SSO integration syncs teams from Okta/Entra with automatic role assignment.
Hierarchical budgets prevent runaway spending. Weighted routing sends 80% of traffic to cost-effective providers with automatic failover to premium options.
MCP tool filtering restricts which tools agents can access. Virtual key permissions ensure agents only call approved models and providers.
Required headers enforce audit trails. RBAC controls who can configure guardrails. Comprehensive logs support SOC 2 Type II, HIPAA, GDPR requirements.
Bind virtual keys to dev/staging/prod API keys. Developers use test keys with lower budgets while production gets dedicated high-limit keys.
[ WHAT'S NEXT ]
Continue with governance, guardrails, MCP, and the rest of the resource library.
PII detection, content moderation, prompt injection defense, and compliance.
SecurityEnforce PII redaction, secrets detection, and prompt injection guardrails on every LLM request and agent tool call from one gateway.
MCPHigh-performance tool execution for AI agents with approvals and audit trails.
[ BIFROST FEATURES ]
Everything you need to run AI in production, from free open source to enterprise-grade features.
01 Governance
SAML support for SSO and Role-based access control and policy enforcement for team collaboration.
02 Adaptive Load Balancing
Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.
03 Cluster Mode
High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.
04 Alerts
Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook and more.
05 Log Exports
Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export capabilities for compliance, monitoring, and analytics.
06 Audit Logs
Comprehensive logging and audit trails for compliance and debugging.
07 Vault Support
Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.
08 VPC Deployment
Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls.
09 Guardrails
Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents.
[ SHIP RELIABLE AI ]
Change just one line of code. Works with OpenAI, Anthropic, Vercel AI SDK, LangChain, and more.
[ FREQUENTLY ASKED QUESTIONS ]
LLM governance is the set of controls that decide who may call which language models, how much they may spend, and what their requests are allowed to do. Bifrost enforces these controls at the AI gateway through virtual keys, budgets, rate limits, and MCP tool filtering, so the rules apply to every application routed through it.
The open-source AI gateway includes virtual keys, hierarchical budgets, rate limits, governance routing, MCP tool filtering, and required headers. Bifrost Enterprise adds RBAC, SSO through OIDC with SCIM provisioning, Access Profiles, signed audit logs of administrative activity, and alerting on budgets and rate limits.
Bifrost checks every budget that applies to a request (provider config, virtual key, team, and customer) and allows it only if all have remaining balance. The cost is deducted from each, so one exhausted budget at any level blocks further requests.
Bifrost returns a 402 when a budget is exhausted and a 429 when a token or request limit is exceeded, with the error type identifying which limit was hit. The key stays active, and requests succeed again once the budget resets or the window expires.
Bifrost separates environments by binding each virtual key's provider configurations to specific provider API key IDs. A development key can use test credentials with a low budget while a production key uses dedicated high-limit credentials.
Bifrost applies the virtual key an agent authenticates with to both its model calls and its MCP tool calls. Tools are denied by default, each MCP client on the key has an explicit allow-list, and the list is enforced again at execution time, as covered in this guide to per-key access control.
LLM governance at the gateway covers traffic routed through Bifrost. For AI tools outside your applications, Bifrost Edge extends the same virtual keys, budgets, and audit logs to desktop apps, browser AI, and coding agents on employee machines, and adds app and MCP server allow and deny decisions per device.