Try Bifrost Enterprise free for 14 days.
Request access
[ BIFROST GOVERNANCE ]

AI Governance Enforced on
Every Request at the Gateway

Control who can access which models, what budget limits apply, and which MCP tools they can use from a single policy layer on the AI gateway.

[ HOW GOVERNANCE WORKS ]

How LLM Governance Works in Bifrost

Each request is authenticated with a virtual key, checked against that key's access rules and every budget above it, and routed only to providers that still have capacity.

Step 01

Authenticate with a virtual key

The caller sends an sk-bf-* key in x-bf-vk, Authorization: Bearer, x-api-key, x-goog-api-key, or api-key. A request with no key is rejected with a 401 when virtual keys are mandatory, and an inactive or expired key returns a 403.

Step 02

Check access rules

Bifrost validates the requested provider and model against the key's provider configurations. A key with no provider configurations blocks all providers, and an empty allowed_models list blocks all models for that provider.

Step 03

Check budgets and rate limits

Every applicable budget must have remaining balance, and the request must pass both the token and request limits at the provider config and virtual key levels. A spent budget returns a 402 and an exhausted rate limit returns a 429.

Step 04

Route and record cost

Traffic is split across allowed providers by weight, skipping any provider over its own limits, and the request's cost is deducted from every budget in the hierarchy.

The configuration below creates a team-scoped virtual key with two providers, a monthly budget, and token and request limits:

Terminal
cURL
curl -X POST http://localhost:8080/api/governance/virtual-keys \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Engineering Team API",
    "provider_configs": [
      { "provider": "openai", "weight": 0.5, "allowed_models": ["gpt-4o-mini"], "key_ids": ["*"] },
      { "provider": "anthropic", "weight": 0.5, "allowed_models": ["claude-3-sonnet-20240229"], "key_ids": ["*"] }
    ],
    "team_id": "team-eng-001",
    "budgets": [{ "max_limit": 100.00, "reset_duration": "1M" }],
    "rate_limit": {
      "token_max_limit": 10000, "token_reset_duration": "1h",
      "request_max_limit": 100, "request_reset_duration": "1m"
    },
    "is_active": true
  }'

[ OSS GOVERNANCE ]

Virtual Keys, Budgets, and Rate Limits

These controls ship in the open-source gateway and are configured through the Web UI, the REST API under /api/governance/*, config.json, or the Helm chart.

Virtual keys

Each key carries model access, budgets, rate limits, and status, and belongs to a user, team, customer, or business unit. Optional expiry rejects the key after a set time without deleting it.

Hierarchical budgets

Independent limits at customer, team, virtual key, and provider levels, with the same cost deducted from each. Reset from one day to one year, with optional UTC calendar alignment.

Token and request rate limits

Cap tokens and requests at the virtual key and provider levels, with 1-minute, 1-hour, or 1-day windows. A provider over its limit is skipped; others on the same key stay available.

Intelligent routing

Split traffic across providers by weight and bind each to specific API key IDs. Separate credentials and spend for development, staging, and production.

Required headers

Reject LLM or MCP requests missing a configured header, such as a tenant or correlation ID, with a 400 before they reach a provider.

[ VIRTUAL KEYS ]

Primary Governance Entity

Virtual keys authenticate requests and enforce access control, budgets, and rate limits per consumer, including securing AI agents with virtual keys and tool filtering.

Model filtering

Restrict which AI models users can access

Provider control

Limit access to specific AI providers

Budget management

Independent cost tracking per virtual key

Rate limiting

Token and request-based throttling

API key restrictions

Bind to specific provider keys

Status control

Instantly enable or disable access

Multiple Header Format Support

x-bf-vk

sk-bf-*

Native Bifrost

Authorization

Bearer

OpenAI-compatible

x-api-key

sk-ant-*

Anthropic-compatible

x-goog-api-key

AI*

Gemini-compatible

[ INTELLIGENT ROUTING ]

Adaptive Load Balancing with Automatic Failover

Distribute traffic across providers with configurable weights and automatic fallback chains.

Adaptive load balancing

Automatically optimizes traffic distribution across providers and keys based on real-time performance metrics.

Automatic failover

Create fallback chains ordered by weight when primary providers fail or hit rate limits

Provider restrictions

Whitelist specific provider-model combinations with empty array defaulting to catalog detection

API key binding

Restrict VKs to specific provider API keys for environment separation (dev/test/prod)

Read the Governance Deep Dive

Virtual keys, hierarchical budgets, weighted routing with automatic failover, and how production teams enforce per-consumer controls without slowing developers.

[ HIERARCHICAL BUDGETS ]

Cost Management at Every Level

Independent cost tracking at Customer, Team, Virtual Key, and Business Unit levels with automatic deduction across all tiers.

Level 1

Customer

Top-level organization with independent budget

Level 2

Team

Department-level budget within customer

Level 3

Virtual Key

Individual access token budget

Level 4

Business Units

Business-unit budget across teams

All Applicable Budgets Must Pass

When a transaction occurs, the same cost deducts from every relevant level simultaneously. A single exhausted budget at any tier blocks the entire request.

[ MCP TOOL FILTERING ]

AI Agent Governance with MCP Tool Filtering

AI agent governance in Bifrost uses one virtual key to define an agent's model access, spend, and MCP tool permissions, so agent traffic counts against the same budgets and rate limits as the team that owns it. Tool connectivity itself is covered on the MCP gateway page.

Per-key tool allow-lists

Tools stay unavailable to a virtual key until an MCP client is configured on it, except clients marked Allow by Default. Each client can allow selected tools, all tools via *, or none.

Headers that can only narrow

Bifrost generates an x-bf-mcp-include-tools header from the key's allow-list. Caller-supplied entries the key does not allow are dropped, and the allow-list is enforced again when the tool executes.

[ ENTERPRISE GOVERNANCE ]

Enterprise Governance with RBAC, SSO, and Audit Logs

Bifrost Enterprise adds identity and administration controls to the open-source policy engine.

CapabilityOpen sourceEnterprise
Virtual keys, budgets, rate limits, routing
MCP tool filtering and required headers
RBAC with system and custom roles
SSO through OIDC and SCIM 2.0 provisioning
Access Profiles and per-user virtual keys
Signed audit logs and budget alerting

Role-based access control

RBAC ships with Admin, Developer, and Viewer system roles plus custom roles, assigned automatically from identity provider groups and claims.

User provisioning

Connect Okta, Microsoft Entra, Keycloak, Zitadel, Google Workspace, Auth0, or any standards-compliant OIDC provider. Teams and business units sync from the IdP; when a user matches several roles, the highest-privilege role applies.

Access Profiles

Reusable policies covering providers, models, budgets, rate limits, and MCP access. Assigning a profile issues each user a write-protected virtual key with isolated counters; profile edits apply on the next request.

Audit logs and alerting

HMAC-signed administrative events that can be archived to S3-compatible storage. Budget and rate-limit alerts go to Slack, Microsoft Teams, PagerDuty, or webhooks.

Deployment options are listed on the Bifrost Enterprise page.

[ COMPLIANCE ]

AICPA SOC
GDPR
ISO 27001
HIPAA

[ ROLE-BASED ACCESS CONTROL ]

Access Control, Flexible to Your Needs

Start with Admin, Developer, and Viewer. Create custom roles as your team structure evolves.

Principle of least privilege

Users receive only the permissions they need for their job function, reducing security vulnerabilities and preventing accidental misconfigurations.

Simplified user management

Assign roles once instead of configuring individual permissions. New team members inherit appropriate access automatically through role assignment.

Audit-ready access tracking

Demonstrate to auditors exactly who has what access. Audit logs track permission changes over time for compliance frameworks like SOC 2 Type II and GDPR.

Custom roles for specialized teams

Create tailored roles for QA teams, security auditors, or compliance officers. Custom roles adapt to your organizational structure.

Three Pre-Configured Roles

Admin

Full control over all Bifrost resources and configurations

Platform engineers, security admins

Developer

Manage technical resources without administrative privileges

Engineering teams, DevOps

Viewer

Read-only access for monitoring and compliance

Finance, compliance, executives

[ CONFIGURATION ]

Four Ways to Configure Governance

Web UI

Visual dashboard for configuring virtual keys, budgets, routing, and RBAC

REST API

Programmatic management via endpoints at /api/governance/*

config.json

Declarative file-based configuration for GitOps workflows

Bifrost CLI

Interactive terminal setup for managing governance settings from your workflow

[ USE CASES ]

Real-World Governance Scenarios

Multi-tenant SaaS platforms

Isolate tenants with virtual keys, enforce per-tenant budgets, and track usage with required headers. Automatic cost allocation across customers.

Enterprise team management

Department-level budgets with team-specific provider access. SSO integration syncs teams from Okta/Entra with automatic role assignment.

Cost control & optimization

Hierarchical budgets prevent runaway spending. Weighted routing sends 80% of traffic to cost-effective providers with automatic failover to premium options.

AI agent security

MCP tool filtering restricts which tools agents can access. Virtual key permissions ensure agents only call approved models and providers.

Regulatory compliance

Required headers enforce audit trails. RBAC controls who can configure guardrails. Comprehensive logs support SOC 2 Type II, HIPAA, GDPR requirements.

Environment separation

Bind virtual keys to dev/staging/prod API keys. Developers use test keys with lower budgets while production gets dedicated high-limit keys.

Ready to Deploy Comprehensive LLM Governance?

Start with OSS governance (virtual keys, budgets, routing) and upgrade to Enterprise when you need RBAC and SSO.

[ BIFROST FEATURES ]

Open Source & Enterprise

Everything you need to run AI in production, from free open source to enterprise-grade features.

01 Governance

SAML support for SSO and Role-based access control and policy enforcement for team collaboration.

02 Adaptive Load Balancing

Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.

03 Cluster Mode

High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.

04 Alerts

Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook and more.

05 Log Exports

Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export capabilities for compliance, monitoring, and analytics.

06 Audit Logs

Comprehensive logging and audit trails for compliance and debugging.

07 Vault Support

Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.

08 VPC Deployment

Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls.

09 Guardrails

Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents.

[ SHIP RELIABLE AI ]

Try Bifrost Enterprise with a 14-day Free Trial

[quick setup]

Drop-in replacement for any AI SDK

Change just one line of code. Works with OpenAI, Anthropic, Vercel AI SDK, LangChain, and more.

1import os
2from anthropic import Anthropic
3
4anthropic = Anthropic(
5 api_key=os.environ.get("ANTHROPIC_API_KEY"),
6 base_url="https://<bifrost_url>/anthropic",
7)
8
9message = anthropic.messages.create(
10 model="claude-3-5-sonnet-20241022",
11 max_tokens=1024,
12 messages=[
13 {"role": "user", "content": "Hello, Claude"}
14 ]
15)
Drop in once, run everywhere.

[ FREQUENTLY ASKED QUESTIONS ]

Common Questions

What is LLM governance?

LLM governance is the set of controls that decide who may call which language models, how much they may spend, and what their requests are allowed to do. Bifrost enforces these controls at the AI gateway through virtual keys, budgets, rate limits, and MCP tool filtering, so the rules apply to every application routed through it.

What is the difference between open source and Enterprise governance in Bifrost?

The open-source AI gateway includes virtual keys, hierarchical budgets, rate limits, governance routing, MCP tool filtering, and required headers. Bifrost Enterprise adds RBAC, SSO through OIDC with SCIM provisioning, Access Profiles, signed audit logs of administrative activity, and alerting on budgets and rate limits.

How do hierarchical budgets work in Bifrost?

Bifrost checks every budget that applies to a request (provider config, virtual key, team, and customer) and allows it only if all have remaining balance. The cost is deducted from each, so one exhausted budget at any level blocks further requests.

What happens when a virtual key exceeds its budget or rate limit?

Bifrost returns a 402 when a budget is exhausted and a 429 when a token or request limit is exceeded, with the error type identifying which limit was hit. The key stays active, and requests succeed again once the budget resets or the window expires.

Can Bifrost separate development and production access?

Bifrost separates environments by binding each virtual key's provider configurations to specific provider API key IDs. A development key can use test credentials with a low budget while a production key uses dedicated high-limit credentials.

How does Bifrost govern AI agents that use MCP tools?

Bifrost applies the virtual key an agent authenticates with to both its model calls and its MCP tool calls. Tools are denied by default, each MCP client on the key has an explicit allow-list, and the list is enforced again at execution time, as covered in this guide to per-key access control.

Does LLM governance cover AI tools on employee laptops?

LLM governance at the gateway covers traffic routed through Bifrost. For AI tools outside your applications, Bifrost Edge extends the same virtual keys, budgets, and audit logs to desktop apps, browser AI, and coding agents on employee machines, and adds app and MCP server allow and deny decisions per device.