Connect Identity with SSO and SCIM Provisioning
- Connect Okta, Entra, Google Workspace, Keycloak, or Zitadel via SSO
- Provision users and assign roles automatically with SCIM
- Sync permissions from IdP groups and claims on every login
One policy layer that keeps AI usage compliant, secure, and accountable with spend controls and audit on every request, without slowing developers down.
[ ENTERPRISE READY: VPC | ON-PREM | AIR-GAPPED ][ OVER 1,000+ TEAMS USE BIFROST ]
[ GOVERNANCE CAPABILITIES ]
One control plane for model access, spend, and compliance across every team and every LLM request.
[ HOW IT WORKS ]
Route every call through a single LLM gateway with access control, budgets, and audit built in.
[ ANY AI REQUEST ]
DENY · over budget / blocked model
request never reaches provider
[ AI GOVERNANCE ON YOUR INFRA ]
Deploy in your own environment with full control over keys, data, and policy
[ UNIFIED GOVERNANCE ]
One stack for access control, spend management, and audit on every model call.
Virtual keys, budgets, rate limits, routing, and audit logs in a single gateway deployment.
Single deploymentEnforce access control, spend caps, and guardrails on every model call before it leaves your perimeter.
Gateway enforcementAttribute cost and usage by team, user, virtual key, model, and provider from one dashboard.
Full attribution[ COMPLIANCE FRAMEWORKS ]
Bifrost Guardrails help organizations meet regulatory requirements with automated detection, redaction, and comprehensive audit trails.




[ FAQ ]
AI governance is the set of policies, controls, and audit processes an organization uses to manage how AI models are accessed, used, and paid for. In practice it covers four things: who can use which models, what data those models can see, how much each team can spend, and whether every request is logged for compliance.
Most governance frameworks stop at policy documents. Enforcing them requires a control point in the request path which is where an LLM gateway comes in.
An AI gateway is a control layer that sits between your applications and every LLM provider. Because all traffic passes through it, it can enforce access rules, budgets, rate limits, and logging on every request, without changing application code.
Bifrost applies identity, access profiles, virtual keys, RBAC, spend limits, and audit logging at this layer, so a policy set once applies across every model, team, and provider.
Without a single control point, each team wires its own provider keys, and the organization loses visibility into who is calling which model and what it costs. That creates shadow AI, unbounded spend, and no audit trail when a regulator asks.
A unified platform consolidates key management, access control, budgets, and logging into one place so governance is enforced by default rather than by policy memo.
Evaluate on five criteria: deployment model (can it run in your VPC, on-prem, or air-gapped?), enforcement depth (does it block policy violations or only report them?), identity integration (SSO, SCIM, existing IdP groups), compliance coverage (audit-log export, PII redaction, framework mapping), and latency overhead in the request path.
Ask specifically whether governance is enforced inline or asynchronously. After-the-fact reporting will not stop a budget overrun or a PII leak.
By tracking token usage per request and attributing it to a team, user, or application, then enforcing hard limits. Bifrost supports hierarchical budgets, per-key and per-team rate limits, and configurable behaviour when a limit is hit: throttle, reject, or fall back to a cheaper model.
Semantic caching reduces spend further by serving repeat requests without a provider call.
Every request through the gateway produces an immutable audit record with user identity, model, token counts, and policy decisions exportable to your SIEM or object storage. PII detection and redaction run inline on request and response bodies. Self-hosted deployment keeps prompts and keys inside your own network boundary, which is often the deciding factor for HIPAA and GDPR workloads.
Budgets cascade from Customer to Team to Virtual Key to Provider. All applicable budgets must pass for a request to proceed. When a transaction occurs, the same cost deducts from every relevant level simultaneously. A single exhausted budget at any tier blocks the entire request.
To know more, read the budgets and rate limits documentation.
Requests return specific HTTP status codes: 402 for budget exceeded, 429 for rate limits exceeded. Virtual keys remain functional for other operations but block LLM requests until budgets reset (based on configured duration) or rate limit windows expire.
See how virtual keys enforce spend and access on every call.
Control who can use which models, cap spend per team, and log every request from one layer. Pair it with AI observability when you need end-to-end traces and cost attribution.
[ BIFROST FEATURES ]
Everything you need to run AI in production, from free open source to enterprise-grade features.
01 Governance
SAML support for SSO and Role-based access control and policy enforcement for team collaboration.
02 Adaptive Load Balancing
Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.
03 Cluster Mode
High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.
04 Alerts
Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook and more.
05 Log Exports
Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export capabilities for compliance, monitoring, and analytics.
06 Audit Logs
Comprehensive logging and audit trails for compliance and debugging.
07 Vault Support
Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.
08 VPC Deployment
Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls.
09 Guardrails
Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents.
[ SHIP RELIABLE AI ]
Change just one line of code. Works with OpenAI, Anthropic, Vercel AI SDK, LangChain, and more.