Try Bifrost Enterprise free for 14 days.
Request access
[ AI OBSERVABILITY ]

One AI Observability Platform
to Monitor and Govern All Your AI

Full visibility into every AI request: content, tokens, cost, latency, routing decisions, and tool calls, exported to the stack you already run.

[ ENTERPRISE READY: VPC | ON-PREM | AIR-GAPPED ]

[ OBSERVABILITY CAPABILITIES ]

LLM and agent observability in production

Track cost and performance per provider, model, and team, with full visibility into every AI request.

Log every request end to end

  • Capture full request and response content, tokens, cost, and latency
  • Query logs by provider, model, status, and time range over the API
  • Turn off content logging when prompts must not be stored
View docs →
Request Log
Live
Provider / ModelTokensCostLatencyStatus
openai / gpt-4o1,204$0.018842msALLOW
anthropic / claude-3.53,880$0.0941.2sALLOW
meta / llama-3-70b620$0.004310msALLOW
openai / gpt-4o-mini2,015$0.006455msALLOW
anthropic / claude-3-haiku990$0.002288msALLOW
openai / gpt-4-turbo5,340$0.1611.6sDENY
mistral / large-21,745$0.021702msALLOW
openai / gpt-4o1,204$0.018842msALLOW
anthropic / claude-3.53,880$0.0941.2sALLOW
meta / llama-3-70b620$0.004310msALLOW
openai / gpt-4o-mini2,015$0.006455msALLOW
anthropic / claude-3-haiku990$0.002288msALLOW
openai / gpt-4-turbo5,340$0.1611.6sDENY
mistral / large-21,745$0.021702msALLOW

Track spend as it happens

  • Meter real cost in USD per provider, model, and virtual key
  • Break spend down by team and customer with governance labels
  • Alert when daily cost crosses a threshold you set
View docs →
Spend Meter
Live
OPENAI · GPT-4O$620
ANTHROPIC · CLAUDE$410
TEAM · GROWTH$780
78%
of daily cap
NEAR LIMIT

Export native Prometheus metrics

  • Scrape success rates, tokens, cost, and cache hits from /metrics
  • Add custom dimensions at runtime with x-bf-dim-* headers
  • Collect asynchronously, with no latency added to requests
View docs →
PER-TOOL METRICS
LIVE · 24h
TOTAL SPEND
$1,284
tool + LLM tokens
P95 LATENCY
142ms
per tool call
TOOL CALLS
38.4K
across 12 servers
LATENCY ms
now
SPEND ATTRIBUTED BY TEAM · CUSTOMERtool · server · vk
team-platform 42%team-growth 31%acme-corp 18%other 9%
STREAM TRACES + METRICS
GrafanaDatadogBigQuery

Trace with OpenTelemetry to any backend

  • Emit OTLP traces using GenAI semantic conventions
  • Send to any OTLP-compatible backend, or your own collector
  • Attach environment, version, and team attributes to every span
View docs →
PER-TOOL METRICS
LIVE · 24h
TOTAL SPEND
$1,284
tool + LLM tokens
P95 LATENCY
142ms
per tool call
TOOL CALLS
38.4K
across 12 servers
LATENCY ms
now
SPEND ATTRIBUTED BY TEAM · CUSTOMERtool · server · vk
team-platform 42%team-growth 31%acme-corp 18%other 9%
STREAM TRACES + METRICS
GrafanaDatadogBigQuery

Debug routing and failover decisions

  • See which routing engine, key, and fallback served each request
  • Track retries and key rotations with the reason for each
  • Watch per-key health flip on failure and recover on success
View docs →
Routing Trace
Live
EngineKeyFallbackReason
weightedkey-a1noneserved OK
prioritykey-c3key-d4429 rate limit
least-latencykey-b2noneserved OK
weightedkey-a1key-e5timeout
prioritykey-d4noneserved OK
least-latencykey-b2key-c3503 unavailable
weightedkey-a1noneserved OK
weightedkey-a1noneserved OK
prioritykey-c3key-d4429 rate limit
least-latencykey-b2noneserved OK
weightedkey-a1key-e5timeout
prioritykey-d4noneserved OK
least-latencykey-b2key-c3503 unavailable
weightedkey-a1noneserved OK

AI agent observability for every MCP tool call

  • Measure duration and failure rate per tool and per server
  • Attribute tool calls to virtual key, team, and customer
  • See MCP tool spend beside LLM token cost
View docs →
AUDIT TRAIL
STREAMING · OTel
TIMECALL · SERVERVIRTUAL KEYLAT
07.412github-mcp · repo.readvk_eng_7f3a82ms
07.588stripe-mcp · charge.createvk_fin_2b1cdeny
08.041postgres-mcp · query.readvk_data_9c45ms
08.517slack-mcp · msg.postvk_ops_4d120ms
09.002filesystem-mcp · fs.readvk_eng_7f3a12ms
09.660jira-mcp · issue.createvk_pm_1a5e210ms
10.128github-mcp · pulls.mergevk_eng_c30f96ms
10.744bigquery-mcp · dataset.listvk_data_9c58ms
07.412github-mcp · repo.readvk_eng_7f3a82ms
07.588stripe-mcp · charge.createvk_fin_2b1cdeny
08.041postgres-mcp · query.readvk_data_9c45ms
08.517slack-mcp · msg.postvk_ops_4d120ms
09.002filesystem-mcp · fs.readvk_eng_7f3a12ms
09.660jira-mcp · issue.createvk_pm_1a5e210ms
10.128github-mcp · pulls.mergevk_eng_c30f96ms
10.744bigquery-mcp · dataset.listvk_data_9c58ms
IMMUTABLE · HASH-CHAINED
SOC 2HIPAAGDPR

[ ON YOUR INFRASTRUCTURE ]

Your telemetry, in your own environment

Bifrost is self-hosted, so observability data is produced and stored where you decide, and exported only to the destinations you configure.

Telemetry stays yours

Logs, traces, and metrics are written by a gateway you run. Nothing is forwarded to a vendor backend unless you configure an exporter yourself.

Export to the stack you already run

Scrape /metrics with Prometheus, ship OTLP traces to Grafana, Datadog, or your own collector, and batch logs to S3, GCS, or BigQuery.

Audit-ready records

Immutable, timestamped request records with virtual key, model, cost, and user, retained on your schedule for SOC 2, HIPAA, and GDPR evidence.

[ COMPLIANCE FRAMEWORKS ]

Built for regulatory compliance

Immutable audit trails and exportable logs across every request, for SOC 2, HIPAA, GDPR, and ISO 27001.

AICPA SOC
GDPR
ISO 27001
HIPAA

[ FAQ ]

Frequently Asked Questions

LLM observability is visibility into what each model call actually did: the request and response content, tokens consumed, real cost, latency, which provider served it, and which tools it invoked. It extends logs, metrics, and traces with AI-specific dimensions that generic APM doesn't capture.

Because every request passes through the gateway, telemetry is captured at that single point rather than instrumented into each application. Bifrost writes logs, emits Prometheus metrics, and produces OTLP traces using GenAI semantic conventions, then exports to whichever backend you configure.

Logs, metrics, and traces, the three classic pillars, plus cost attribution, which is specific to AI. Token spend has to be traceable to a provider, model, team, and virtual key, or you have performance data with no way to tie it to budget.

Yes. Bifrost emits OTLP traces to any OTLP-compatible backend: Grafana, Datadog, New Relic, Honeycomb, or a collector you run. Prometheus scrapes /metrics natively. Request logs export to S3, GCS, BigQuery, Kafka, or Pub/Sub on a schedule.

No. Observability adds 0ms on the request path. Metrics, logs, and traces are collected asynchronously off the hot path: Prometheus counters update in memory, and OTLP spans are batched and flushed by a background exporter.

Yes. Content logging is a separate switch from metering. Turn it off and Bifrost still records tokens, cost, latency, model, status, and governance attribution, without persisting prompts or completions.

Bifrost meters real cost in USD per request using per-model pricing for input, output, cached, and reasoning tokens, then attributes it to the provider, model, virtual key, team, and customer that produced it.

Yes. Send x-bf-dim-* headers on a request and Bifrost attaches them as labels at runtime, so you can slice metrics by environment, version, tenant, or any other dimension you care about.

Yes. Tool calls are logged with duration, failure rate, server, and tool name, attributed to the same virtual key and team as the parent LLM request, so tool spend sits beside token spend. See the MCP gateway for how tool traffic is governed at the same layer.

See every AI request your organization makes

Track cost and performance per provider, model, and team, on infrastructure you run yourself.

[ BIFROST FEATURES ]

Open Source & Enterprise

Everything you need to run AI in production, from free open source to enterprise-grade features.

01 Governance

SAML support for SSO and Role-based access control and policy enforcement for team collaboration.

02 Adaptive Load Balancing

Automatically optimizes traffic distribution across provider keys and models based on real-time performance metrics.

03 Cluster Mode

High availability deployment with automatic failover and load balancing. Peer-to-peer clustering where every instance is equal.

04 Alerts

Real-time notifications for budget limits, failures, and performance issues on Email, Slack, PagerDuty, Teams, Webhook and more.

05 Log Exports

Export and analyze request logs, traces, and telemetry data from Bifrost with enterprise-grade data export capabilities for compliance, monitoring, and analytics.

06 Audit Logs

Comprehensive logging and audit trails for compliance and debugging.

07 Vault Support

Secure API key management with HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, and Azure Key Vault integration.

08 VPC Deployment

Deploy Bifrost within your private cloud infrastructure with VPC isolation, custom networking, and enhanced security controls.

09 Guardrails

Automatically detect and block unsafe model outputs with real-time policy enforcement and content moderation across all agents.

[ SHIP RELIABLE AI ]

Try Bifrost Enterprise with a 14-day Free Trial

[quick setup]

Drop-in replacement for any AI SDK

Change just one line of code. Works with OpenAI, Anthropic, Vercel AI SDK, LangChain, and more.

1import os
2from anthropic import Anthropic
3
4anthropic = Anthropic(
5 api_key=os.environ.get("ANTHROPIC_API_KEY"),
6 base_url="https://<bifrost_url>/anthropic",
7)
8
9message = anthropic.messages.create(
10 model="claude-3-5-sonnet-20241022",
11 max_tokens=1024,
12 messages=[
13 {"role": "user", "content": "Hello, Claude"}
14 ]
15)
Drop in once, run everywhere.