Bifrost AI Gateway: Monitor Every Request, Token, and Cost
A guide to AI gateway observability, covering the telemetry captured per request and how to export it to Prometheus, OpenTelemetry or Datadog.
TL;DR
- AI spend running across several providers cannot be attributed from provider billing dashboards, because none of them share a schema, a schedule, or a consumer identifier.
- AI gateway observability closes that gap by recording provider, model, virtual key, token counts, latency, cost, and status for every request in one consistent format, with no instrumentation in application code.
- Bifrost captures that telemetry while adding 11 microseconds of overhead per request at 5,000 requests per second in sustained benchmarks.
- Metrics leave the gateway in formats existing stacks already consume: a Prometheus endpoint, OTLP traces and metrics, and a native Datadog connector.
- Request logs and audit logs are different records. Request logs cover the traffic; audit logs cover administrative changes to the configuration.
When organizations run AI workloads across OpenAI, Anthropic, AWS Bedrock, and Google Vertex simultaneously, understanding total spend, per-consumer usage, and provider error rates requires data from every request passing through the system. Without a centralized gateway, teams fall back on per-provider billing dashboards that cannot be correlated across vendors or attributed to individual teams and applications. Bifrost, an open-source AI gateway built in Go, solves this by routing all LLM traffic through a single layer that captures structured telemetry for every call, regardless of which provider handles it. That single layer is what makes AI gateway observability possible without touching application code.
What AI Gateway Observability Covers
AI gateway observability is visibility into the full lifecycle of every LLM request passing through the gateway: the provider called, the model used, the token count, the latency, the cost, the virtual key (consumer identity), and any errors or fallbacks triggered. It allows teams to attribute AI spend, detect anomalies, and trace failures without manually aggregating per-provider dashboards.
What it does not cover matters too. Gateway observability sees the request and the response, not the reasoning inside the model. It answers what was called, by whom, at what cost, and whether it succeeded. Judging whether the output was correct is a separate discipline with separate tooling. What to measure at the gateway, and where draws that line in more detail.
The Observability Gap in Direct Provider API Access
When application code calls provider APIs directly, each provider returns its own billing data in its own format on its own schedule. OpenAI usage appears in the OpenAI dashboard, Anthropic usage in the Anthropic console, Bedrock usage in AWS Cost Explorer. None of these sources share a common schema, so correlating a spike in total AI spend to a specific provider, model, or application requires manual cross-referencing that rarely happens in real time.
The attribution problem compounds when multiple teams share the same provider account. Without request-level tagging enforced at the gateway, there is no reliable way to determine which team consumed which tokens. Ad hoc solutions, like adding custom headers or logging request metadata in application code, are inconsistent across codebases and break whenever a new model or provider is added.
Error correlation is equally fragmented. A 503 from AWS Bedrock and a 429 from OpenAI both result in a failed request for the end user, but diagnosing the root cause means checking two separate dashboards with different retention windows and log formats. Real-time error rate visibility across all providers simply does not exist without a shared routing layer sitting in front of all provider traffic. The same fragmentation is why managing LLM spend across providers is usually the first thing teams centralize.
Bifrost's Built-In Observability Layer
The Bifrost AI gateway captures structured telemetry for every request at the gateway level, adding only 11 microseconds of overhead at 5,000 requests per second, so the observability layer does not become a bottleneck in high-throughput production systems.
Every request log includes the provider called, the model selected, the virtual key that authenticated the request, the prompt and completion token counts, the end-to-end latency, the calculated cost, and the response status. This gives a unified record across all providers in a consistent schema without any instrumentation in application code.
Per-virtual-key cost and token usage is available from the moment a key is created. Because every request carries the virtual key identity, the dashboard can break down total spend, daily token consumption, and request volume by consumer without any post-processing. A team that owns a specific virtual key can see exactly what it has spent across all providers it is authorized to reach.
Provider error rates are tracked continuously. When a provider returns a 5xx or a rate limit error and automatic fallback kicks in, that event is recorded against the originating virtual key and the failing provider, so teams can see both the user-facing impact and the provider-level instability over time. Semantic cache hit rates are surfaced alongside request counts, showing exactly how much token spend cache hits are avoiding. Rate limit consumption per key is tracked in real time against the configured budgets and rate limits, so teams can see how close any consumer is to hitting its ceiling before requests start failing.
Exporting Metrics to Existing Observability Stacks
Exporting gateway telemetry does not mean adopting a new monitoring platform. Bifrost as the routing layer emits metrics and traces in the formats below, so AI traffic lands in the dashboards and alerting rules a team already operates rather than in a separate console.
The Prometheus metrics endpoint is available by default, though once Bifrost authentication is enabled the scraper needs credentials, and without them Prometheus receives 401s and scraping fails silently. Configure your Prometheus scrape interval (every 15 or 30 seconds is typical) and all request counts, token totals, latency histograms, error rates, and cache hit rates flow into existing Grafana dashboards or any other Prometheus-compatible visualization layer. Alert rules can be written against these metrics using standard PromQL without any vendor-specific query language.
OpenTelemetry / OTLP export sends traces and metrics over the OTLP protocol to Grafana Cloud, Honeycomb, New Relic, Datadog, or any OTLP-compatible collector, self-hosted included. Each LLM request becomes a trace span with all relevant attributes attached, so AI calls appear in the same distributed tracing view as the rest of the application stack. The OpenTelemetry GenAI semantic conventions define the standard attribute names for those spans, which is what keeps the telemetry portable between backends. A slow LLM response shows up in the same waterfall as the database query and the API call it is paired with.
The Datadog plugin maps Bifrost's request telemetry to Datadog's LLM Observability schema and APM trace format. Token usage and cost attribution land in the LLM Observability dashboard without any custom mapping, and the connector links LLM spans to the broader APM traces they belong to so teams can see the full request context.
Log exports offload request and response payloads to object storage while the logs database keeps searchable metadata, indexes, and pointers. S3 and GCS are the supported destinations today; Azure Blob, local filesystem, and direct data-warehouse targets are not implemented. Teams that need long retention for audit purposes, or that want to run cost attribution queries against historical traffic, can query the exported objects from their own data lake without additional middleware.
Tracing AI Spend Per Team, Application, and Model
Virtual keys are the mechanism that makes spend attribution work without any changes to application code. Each consumer, whether a team, a specific application, or an individual user, gets its own virtual key. All requests made with that key are tagged with the key identity at the gateway layer, so every token count and cost figure in the observability data carries a consumer identifier.
The result is that per-team and per-application spend dashboards are available without requiring each team to implement its own logging or tagging logic. When a new model is added or a new provider is onboarded, the attribution still works because the virtual key is validated and recorded by the gateway before the request is forwarded to any provider.
Model-level breakdowns follow from the same data. Because the gateway records the exact model used on every call, it is straightforward to see which models drive the most token spend, which are called most frequently, and which carry the highest per-request cost. Comparisons across providers for equivalent models are available in a single view without cross-referencing two billing consoles. Tracking GenAI spend through an enterprise gateway covers how teams structure keys so the attribution stays meaningful as the organization grows, and cost tracking for coding agents covers the case where the consumer is a developer tool rather than a service.
Alerting on AI Spend and Error Rates
Collecting observability data is only useful if it drives action before problems compound. The metrics Bifrost exports map directly to alert conditions that matter for production AI systems.
Per-key budget limits prevent overspend before it happens. When a virtual key approaches its monthly or daily spend budget, or its configured token rate limit, the gateway enforces the limit rather than passing the request to the provider.
Enforcement alone is a blunt signal, because the first sign of trouble is a rejected request. Alerting in Bifrost Enterprise closes that gap natively: alert rules are CEL expressions evaluated against live governance metrics, budgets and rate limits, on a 60-second sweep, scoped to a virtual key, team, or customer. When a threshold is crossed, notifications dispatch to Slack, Microsoft Teams, PagerDuty, or any HTTP webhook, with per-rule cooldowns to prevent alert storms and a recorded history of every evaluation and delivery attempt. Budget state is not exposed as a Prometheus metric, so this is the path for spend alerting rather than a PromQL rule.
Rate limit error spikes (HTTP 429s from providers) indicate provider-side throttling that may not yet be visible in provider dashboards. Alerting on the rate limit error rate from Bifrost's Prometheus metrics catches throttling events in real time. Fallback activation frequency is a related signal: a sudden increase in the number of requests that trigger automatic fallback indicates provider instability, even if the provider's own status page has not been updated.
P99 latency anomalies are detectable from Bifrost's latency histograms. A provider-side degradation that doubles median response time will appear in the latency percentile metrics minutes before it surfaces in user-facing error rates. Datadog monitors, Prometheus alerting rules, and OTLP-based alerting in Honeycomb or New Relic all consume these metrics natively, so teams can set up latency and error rate alerts using the same tooling they already use for the rest of the stack. Budget tracking and spend alerts in production compares how different gateways expose these signals.
Compliance Logging for AI Requests: Request Logs vs Audit Logs
Compliance programs need two different records, and Bifrost keeps them separately. Confusing the two is the most common mistake when scoping an audit, because each answers a question the other cannot.
| Record | What it captures | Where it lives |
|---|---|---|
| Request logs | Every call through the gateway: virtual key, provider, model, timestamp, token counts, latency, cost, status | The logs database, with payloads optionally offloaded to object storage |
| Audit logs | Administrative activity: who changed what, when, and which resource was affected | Enterprise audit store, optionally HMAC-signed and archived |
Request logs are the record of a specific AI call: they support SOC 2 Type II evidence collection, HIPAA access logging requirements, ISO 27001 operational records, and GDPR data processing records without requiring any additional instrumentation in application code. Audit logs answer the separate question of who altered a virtual key, a budget, or a routing rule, and their entries can be HMAC-signed so a reviewer can verify they were not tampered with, retained for a configurable number of days, and archived to object storage.
For organizations with long-term retention requirements, both records reach object storage: request payloads through log exports and audit events through audit-log archival, in either case to S3 or GCS. The export schema is consistent and documented, so compliance and security teams can write queries against historical request data without needing to understand provider-specific log formats. For agent workloads specifically, auditing every AI tool call at the MCP gateway extends the same record to tool execution.
Frequently Asked Questions About AI Gateway Observability
Does gateway observability replace application-level tracing?
No, it completes it. The gateway sees the provider call, not the application logic around it. Because Bifrost emits each request as an OTLP span, those calls join the same distributed trace as the database queries and service hops they sit between, so a slow model response appears in the same waterfall rather than in a separate tool.
How much latency does the observability layer add?
Very little. Bifrost adds 11 microseconds of overhead per request at 5,000 requests per second in sustained benchmarks, which is negligible against a model call measured in hundreds of milliseconds. Logging runs asynchronously, so capturing telemetry does not sit in the response path.
Can request logs be exported straight to a data warehouse?
Not directly. S3 and GCS are the supported export destinations; Azure Blob, local filesystem, and direct warehouse targets are not implemented. The usual pattern is to export to object storage and query those objects from the warehouse or lake that already reads that bucket, keeping the gateway out of the ingestion path.
Do you need both Prometheus and OpenTelemetry, or just one?
One is usually enough, and the choice follows the receiving stack. Prometheus suits numeric time series and alerting rules written in PromQL. OTLP suits distributed tracing, where each request becomes a span with attributes. Teams running both typically scrape metrics with Prometheus and send traces over OTLP.
How is spend attributed when several teams share one provider account?
Through virtual keys. Each team, application, or user authenticates with its own key, and the gateway records that identity on every request before forwarding it. Attribution therefore happens at the routing layer rather than in application code, and it keeps working when a new model or provider is added.
Does the gateway record prompts and completions themselves?
Yes, and where they are stored is configurable. Request and response payloads can be offloaded to object storage while the logs database keeps searchable metadata, indexes, and pointers. That keeps the database small while leaving full payloads available for retention or analysis in your own storage.
Get Full Visibility with Bifrost
Organizations running AI in production across multiple providers need AI gateway observability: request-level visibility, per-consumer attribution, and real-time alerting. Bifrost provides all of this from a single gateway, with export to Prometheus, OpenTelemetry, Datadog, and S3-compatible storage, so teams can see exactly what their AI systems are doing without building custom aggregation pipelines.
The practical sequence is short. Route one application through the gateway, give it its own virtual key so its spend is separable from the start, point the Prometheus endpoint or OTLP exporter at the stack already in use, then set a budget and an error-rate alert before adding the second application. Deciding what to measure at the gateway is worth settling before the second team onboards.
To see how Bifrost fits into your observability stack, schedule a demo with the team, or explore the benchmarks and governance resources to understand the full scope of what the gateway covers.