---
description: "See what an LLM gateway does on every request: authentication, budgets, routing, failover, guardrails, caching, and telemetry, and how to deploy it."
title: "What an LLM Gateway Actually Does: A Guide for AI Infrastructure Teams"
image: https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastruct-bifrost-paper-waves.optimized.png
---

Try Bifrost Enterprise free for 14 days. [Request access](https://www.getmaxim.ai/#enterprise-trial)

**TL;DR**

- An [LLM gateway](https://www.getmaxim.ai/llm-gateway) runs a fixed pipeline on every model call: authenticate the caller, check governance, route, apply guardrails, forward, handle the response, and record telemetry.
- Applications hold gateway-issued virtual keys while provider API keys stay inside the gateway, so rotating or revoking a credential is a single change.
- Bifrost enforces budgets at four levels (customer, team, virtual key, and provider configuration) and rate limits on both requests and tokens.
- Bifrost adds 11 microseconds of overhead per request at 5,000 RPS; optional features such as semantic caching and external guardrail providers add their own network calls.
- The same virtual key governs MCP tool calls, so an agent's model access and tool access are set in one place.

**An LLM gateway is a control layer purpose-built for model API traffic.** [**Bifrost**](https://www.getmaxim.ai/bifrost) **is an open-source LLM gateway built for sub-millisecond overhead, full governance, and no vendor lock-in.**

When an application sends a request to an LLM provider, it passes through a chain of concerns that have no equivalent in standard API infrastructure: the token economy affects cost at every hop, streaming responses hold connections open for seconds or minutes, provider rate limits are per-key and per-minute rather than per-IP, and the content of a prompt is a security surface that HTTP middleware was never designed to inspect. An LLM gateway handles all of these at the infrastructure layer, before a single token is billed. [Bifrost](https://www.getmaxim.ai/bifrost), the [high-performance open-source LLM gateway](https://github.com/maximhq/bifrost) built in Go by Maxim AI, is built for teams that need these capabilities with [11 microseconds of added overhead at 5,000 RPS](https://www.getmaxim.ai/bifrost/resources/benchmarks) and no external service dependency.

## What an LLM Gateway Does on Every Request

An LLM gateway intercepts every model API call your application makes and runs it through a fixed sequence of identity, policy, routing, and logging steps before and after the provider call. On each request, it executes a deterministic pipeline in this order:

1. **Authenticate the caller** using a virtual key, API key header, or bearer token
2. **Check governance rules:** budget remaining, rate limit headroom, model allowlist membership
3. **Select a provider and model** based on routing rules, weights, and fallback chains
4. **Apply input policies:** content guardrails, PII detection, prompt injection scanning
5. **Forward the request** to the upstream provider using the gateway-held credential
6. **Handle the response:** streaming pass-through, output policy checks, semantic cache population
7. **Write telemetry:** token counts, latency, cost, policy decisions, caller identity

Bifrost evaluates authentication, governance, and routing in-process from in-memory state, so the only mandatory network call is the model call itself; optional features such as an external guardrail provider or a semantic cache add their own round-trips. The caller's application code sees a single OpenAI-compatible endpoint. Provider API keys never leave the gateway, and the application only holds a virtual key. The full [request flow documentation](https://docs.getbifrost.ai/architecture/core/request-flow) covers how each stage is implemented at the architecture level.

![A request moves left to right through virtual key authentication, governance checks, routing, input guardrails, and the provider call, with early exits for auth, budget, and guardrail failures](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/llm-gateway-request-pipeline.png)

*Figure 1: Cheap identity and policy checks run before any token is billed, so a rejected request never reaches a provider.*

The [Gartner Market Guide for AI Gateways 2025](https://www.gartner.com/en/documents/7051698) projects that 70% of software engineering teams building multimodel applications will use AI gateways by 2028, up from 25% in 2025. The driver is that each step in the pipeline above represents a class of production failures that accumulates quickly without a centralized enforcement point. The [LLM gateway deep dive](https://www.getmaxim.ai/articles/what-is-an-llm-gateway-a-deep-dive-into-the-backbone-of-scalable-ai-applications/) covers why this layer becomes the backbone of scalable AI applications.

## How an LLM Gateway Differs from an API Gateway

An LLM gateway differs from a traditional API gateway in what it meters and inspects. An API gateway routes and rate-limits HTTP requests to backends it does not understand; an LLM gateway meters tokens and dollars, holds model provider credentials, fails over between providers, and inspects prompt and completion content.

| Dimension | Traditional API gateway | LLM gateway |
| --- | --- | --- |
| Rate limiting unit | Requests per client | Requests and tokens per virtual key |
| Cost tracking | Not model-aware | Dollar spend per key, team, and customer, priced per model |
| Upstream failure | Retry or error for one backend | Retries, key rotation, and fallback to another provider |
| Payload awareness | Headers and paths | Prompts, completions, and tool calls |
| Credentials | Client auth to owned services | Holds provider API keys; clients use virtual keys |

An existing API gateway can sit in front of an LLM gateway, but it does not replace the token accounting, provider failover, and content policy that live in this layer. The [AI gateway architecture overview](https://www.getmaxim.ai/articles/what-is-an-ai-gateway-architecture-features-and-why-it-matters/) shows where the two layers meet.

## The Request Lifecycle in Detail

Each stage of the LLM gateway request lifecycle maps to a configurable Bifrost feature, covered below in request order.

### Authentication and Credential Isolation

The LLM gateway holds all provider API keys. Applications authenticate using [virtual keys](https://docs.getbifrost.ai/features/governance/virtual-keys), gateway-issued credentials that carry scoped permissions but no provider secrets. This eliminates credential sprawl: rotating a compromised provider key means updating one record in the gateway, not hunting down environment variables across services. Revoking an application's access means deactivating its virtual key.

![Two services authenticate to Bifrost with virtual keys, and Bifrost calls OpenAI and Anthropic with provider API keys that never leave the gateway](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/virtual-key-credential-isolation.png)

*Figure 2: Rotating a provider key or revoking an application is one change in the gateway, not a hunt through service configs.*

Virtual keys also carry the governance context for the request: which providers and models the caller can access, what budget remains, and what rate limits apply.

### Routing, Failover, and Load Balancing

Once authenticated, the gateway applies routing logic. Bifrost supports weighted routing across multiple provider configurations within a single virtual key: 60% of traffic to OpenAI, 40% to a Bedrock endpoint, with automatic failover to the Bedrock path if the OpenAI call returns 5xx errors or times out.

[Load balancing across API keys](https://docs.getbifrost.ai/features/keys-management) for the same provider distributes request volume across the provider's per-key rate limit buckets, pooling the throughput of keys that belong to separate provider accounts or projects without any application-layer logic. [Automatic fallbacks](https://docs.getbifrost.ai/features/fallbacks) apply exponential backoff and retry sequencing, including cross-provider failover where a failure on one provider routes to a pre-configured backup on a different provider.

This is the layer that makes multi-provider deployments operationally viable. Without it, every provider failure requires application-level handling, which means every team that calls an LLM must re-implement the same retry and fallback logic.

### Governance: Budgets, Rate Limits, and Model Allowlists

The gateway is the only layer in the stack that sees every LLM call from every application. It is therefore the only layer that can enforce organization-wide cost and access policy reliably.

[Budget enforcement](https://docs.getbifrost.ai/features/governance/budget-and-limits) runs at four levels: the provider configuration inside a virtual key, the virtual key (per-application or per-developer), the team (aggregate for a department), and the customer (organization-wide cap). All applicable budgets are checked independently on each request. When a virtual key, team, or customer budget is exhausted, the gateway returns HTTP 402 before the request reaches a provider; a provider configuration that has spent its budget is excluded from routing instead. Budgets reset on configurable windows (daily, weekly, monthly, quarterly, or yearly) with calendar-aligned or rolling options, as detailed in [hierarchical LLM budget management](https://www.getmaxim.ai/articles/llm-budget-management-virtual-keys-and-hierarchical-spend-controls/).

Rate limiting runs on two independent dimensions, request frequency and token throughput, at both the virtual key and provider configuration levels. A virtual key can carry 1,000 requests per minute alongside 2 million tokens per hour, with each dimension tracked and enforced separately. This matters for agentic workloads where a single session can generate many requests with small prompts, or few requests with very large context windows. For a full breakdown of how these dimensions compare across [LLM gateway](https://www.getmaxim.ai/articles/top-5-llm-gateways-in-2026-a-production-ready-comparison/) implementations, the [AI gateway buyer's guide](https://www.getmaxim.ai/bifrost/resources/buyers-guide) covers governance depth as a standalone evaluation criterion.

Model allowlists attach to each provider configuration within a virtual key: `"allowed_models": ["claude-sonnet-4-6", "claude-haiku-4-5-20251001"]`. A request asking for a model not in the allowlist returns HTTP 403 before any token is consumed. Platform teams use this to enforce tiered access: lower-cost models for high-volume batch jobs, premium models only for approved workflows.

### Semantic Caching

LLM API costs compound on repeated similar queries. A customer support application that answers the same category of questions hundreds of times per day sends hundreds of API calls where a small cache would serve most of them.

[Semantic caching](https://docs.getbifrost.ai/features/semantic-caching) stores responses by vector embedding and returns cached results for queries that fall within a configurable similarity threshold of a prior query. Standard HTTP caches operate on exact string match; semantic caching matches on meaning, which recovers significant cost savings in conversational and Q&A workloads. Bifrost checks an exact-match hash first, which costs one vector store round-trip, and runs the [semantic search for similar LLM queries](https://www.getmaxim.ai/articles/how-to-optimize-llm-cost-and-latency-with-semantic-caching/) only on a miss, which adds an embedding call.

### Content Guardrails

[OWASP ranks prompt injection as the top security risk for LLM applications](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) as of 2025, and sensitive information disclosure as the second. Both risks live in the content of prompts and completions, a layer that standard API gateways have no visibility into.

Bifrost's [guardrails](https://docs.getbifrost.ai/enterprise/guardrails) validate both prompt inputs and model outputs inline, before they transit the gateway in either direction. Organizations configure rules using CEL expressions and attach profiles from three Bifrost-managed providers (Prompt Guardrails, Custom Regex, and Secrets Detection) or from external providers including AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, Gray Swan Cygnal, and Patronus AI. Custom Regex and Secrets Detection run in-process with no external service call, adding no round-trip latency. Violations return HTTP 446 (blocked) or 246 (passed with a logged warning) with structured violation detail for downstream handling.

### Observability

[Menlo Ventures estimated](https://menlovc.com/perspective/2025-mid-year-llm-market-update/) that enterprise spend on model APIs rose from $3.5 billion in November 2024 to $8.4 billion by mid-2025, and many teams have limited visibility into how that spend is distributed across applications, teams, and model choices. The gateway is the only layer that sees every call, making it the natural source of truth for AI infrastructure observability.

Bifrost captures per-request telemetry (provider, model, token counts at input and output, latency, cost, caller identity, and policy outcomes) without any instrumentation in application code. This telemetry feeds native [Prometheus metrics](https://docs.getbifrost.ai/features/observability/prometheus) and [OpenTelemetry traces](https://docs.getbifrost.ai/features/observability/otel) directly to Grafana, Datadog, New Relic, and any OTLP-compatible collector. Every request produces a structured log entry regardless of whether the application developer added any logging. The [gateway-level LLM observability metrics](https://www.getmaxim.ai/articles/llm-observability-at-the-gateway-what-to-measure-and-where/) worth alerting on are covered separately.

## MCP Gateway and Agent Traffic

An LLM gateway that also acts as an MCP gateway applies its identity, access, and logging controls to agent tool calls as well. In agentic architectures, a single user request triggers multiple LLM calls interleaved with tool invocations through the [Model Context Protocol](https://modelcontextprotocol.io/). Each tool call is an execution surface that requires its own authentication, access control, and audit record.

Bifrost extends the same governance model to tool traffic: it acts as both an MCP client (connecting to external tool servers on behalf of agents) and an MCP server (exposing configured tools to MCP-compatible clients). [Tool filtering per virtual key](https://docs.getbifrost.ai/features/governance/mcp-tools) applies a deny-by-default allowlist to every tool call. LLM calls and MCP tool calls are recorded in the same [request logs](https://docs.getbifrost.ai/features/observability/default), giving infrastructure teams a unified view of what each agent session actually consumed. The agent-side design is covered in [what an MCP gateway does for production AI agents](https://www.getmaxim.ai/articles/what-is-an-mcp-gateway-a-guide-for-production-ai-agents/).

![An agent session sends a model call, an MCP tool call, and a second model call through Bifrost, which checks the tool allowlist and logs all three](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/agent-traffic-one-log.png)

*Figure 3: Tokens and tool calls from one agent session land in the same logs under the same virtual key.*

## Deployment Considerations for Infrastructure Teams

An LLM gateway sits in the critical path of every model call, which means its operational properties matter directly: the latency it adds, where its governance state lives, where request data is processed, and how much code must change to adopt it.

**Latency overhead** is the first consideration. A gateway that adds 200ms per call compounds across multi-step agent chains. Bifrost [benchmarks at 11 microseconds of added latency at 5,000 RPS](https://www.getmaxim.ai/bifrost/resources/benchmarks) in sustained load tests on standard cloud instances, making it a transparent layer for even latency-sensitive workloads.

**State management** affects correctness and scalability. Bifrost holds all governance state in memory (provider configs, virtual key permissions, budget counters, rate limit state) for sub-millisecond policy evaluation without database round-trips. The [OSS build](https://docs.getbifrost.ai/overview) handles approximately 3,000 to 5,000 RPS on a single instance with a Postgres backend for persistence. [Bifrost Enterprise clustering](https://docs.getbifrost.ai/enterprise/clustering) uses gossip-based membership and a dedicated gRPC channel that synchronizes governance counters, virtual keys, and routing rules across nodes, enabling horizontal scaling with consistent enforcement.

![A load balancer spreads application traffic across three Bifrost nodes that sync membership over gossip and budget counters over gRPC, and each node calls model providers](https://articles-images-cdn.t3.tigrisfiles.io/diagrams/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/bifrost-cluster-in-vpc.png)

*Figure 4: Every node enforces the same budgets and rate limits because counters replicate across the cluster.*

**Deployment model** determines the data perimeter. For regulated industries and teams with data residency requirements, [in-VPC deployment](https://docs.getbifrost.ai/enterprise/invpc-deployments) keeps all request bodies, telemetry, and audit logs within the organization's private network. A hosted gateway, by design, processes request data outside that perimeter.

**Drop-in migration** determines adoption friction. Bifrost exposes an OpenAI-compatible API on a single endpoint. Migrating an existing application that uses the OpenAI SDK requires changing the base URL and, when governance is enabled, replacing the provider API key with a Bifrost virtual key in the same field. No application code changes, no SDK swaps. The [drop-in replacement guide](https://docs.getbifrost.ai/features/drop-in-replacement) covers migration for OpenAI, Anthropic, Bedrock, and Google GenAI SDK users.

## Evaluating an LLM Gateway

An LLM gateway evaluation should measure six properties under production conditions: latency overhead, governance depth, agent and MCP support, deployment options, compliance evidence, and migration effort. The enterprise buying criteria are laid out in [the complete LLM gateway guide for enterprise AI](https://www.getmaxim.ai/articles/what-is-an-llm-gateway-complete-guide-for-enterprise-ai-in-2026/). Infrastructure teams evaluating LLM gateways should assess on these dimensions before selecting:

- **Latency overhead at production RPS:** not demo load, sustained benchmark
- **Governance depth:** hierarchical budgets, token and request rate limits, model allowlists, per-credential access control
- **MCP and agent support:** tool-level access control, unified request logs across LLM and tool calls
- **Deployment options:** self-hosted, in-VPC, on-premises, air-gapped
- **Compliance evidence:** immutable audit logs, SOC 2, GDPR, HIPAA
- **Migration path:** drop-in SDK compatibility vs. required code changes

The [LLM Gateway Buyer's Guide](https://www.getmaxim.ai/bifrost/resources/buyers-guide) maps each of these dimensions to a structured evaluation framework with specific questions for each capability area. For teams evaluating Bifrost against specific infrastructure requirements, the [resources hub](https://www.getmaxim.ai/bifrost/resources) covers benchmarks, governance, and the [MCP gateway](https://www.getmaxim.ai/mcp-gateway) in full detail. Scalability testing is covered in [how to evaluate an LLM gateway for enterprise scalability](https://www.getmaxim.ai/articles/how-to-evaluate-an-llm-gateway-for-enterprise-scalability/).

## LLM Gateway FAQs

### What does an LLM gateway do on each request?

An LLM gateway authenticates the caller, checks budgets, rate limits, and model permissions, selects a provider and API key, applies input guardrails, forwards the call, applies output checks, and records tokens, cost, and latency. Bifrost runs these steps in a fixed order behind one OpenAI-compatible endpoint, so application code sees a single API regardless of which provider serves the request. The [architecture behind an LLM gateway](https://www.getmaxim.ai/articles/what-is-an-llm-gateway-a-deep-dive-into-the-backbone-of-scalable-ai-applications/) is covered separately.

### Does an LLM gateway replace an API gateway?

No. An API gateway handles general HTTP concerns such as ingress, TLS, and routing to internal services. An LLM gateway handles concerns specific to model traffic: token-based rate limits, per-model cost tracking, provider failover, and prompt content policy. Many teams run both, with the API gateway in front of the LLM gateway.

### Where are provider API keys stored in an LLM gateway?

Provider API keys are stored in the gateway, not in applications. Applications authenticate with Bifrost virtual keys, which carry permissions and budgets but no provider secrets. Enterprise deployments can load provider keys from AWS Secrets Manager, GCP Secret Manager, or HashiCorp Vault through [secret management](https://docs.getbifrost.ai/enterprise/secret-management), so rotating a provider key never requires an application redeploy.

### Does semantic caching always reduce latency?

No. A direct hash hit is served after one vector store lookup and is much faster than a model call. A semantic lookup must first embed the incoming request, so a semantic hit costs roughly one embedding round-trip, and a semantic miss adds that embedding call on top of the full model call. Semantic caching pays off on workloads with many repeated, similarly worded questions.

### How does an LLM gateway scale across multiple instances?

A single Bifrost instance holds governance state in memory and handles roughly 3,000 to 5,000 RPS. Beyond that, Bifrost Enterprise clusters nodes: membership runs over gossip, and budget counters, virtual keys, and routing rules replicate over gRPC, so a budget or rate limit means the same thing on every node behind the load balancer.

## Getting Started

Bifrost deploys as a Docker container or binary. The [gateway setup guide](https://docs.getbifrost.ai/quickstart/gateway/setting-up) covers first provider configuration, virtual key creation, and a first authenticated request. For enterprise requirements including clustering, RBAC, SSO integration, and in-VPC deployment, [Bifrost Enterprise](https://www.getmaxim.ai/bifrost/enterprise) is available with a 14-day trial. Once traffic flows, the order in which routing rules, fallbacks, and budgets apply is covered in [LLM gateway routing, fallback, and governance in Bifrost](https://www.getmaxim.ai/articles/llm-gateway-routing-fallback-and-governance-in-bifrost/).

To see how Bifrost fits into an existing AI infrastructure stack, [book a demo](https://getmaxim.ai/bifrost/book-a-demo) with the Bifrost team.

## Read next

[![Top 5 AI Gateways for Controlling Shadow AI in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-controlling-shadow-ai-bifrost-isometric.png) Shadow AI is the use of AI tools, models, and MCP servers that security teams have not approved and cannot see. This guide ranks five AI gateways for controlling it, including Bifrost with Bifrost Edge, Kong AI Gateway, Cloudflare AI Gateway, and Gravitee.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-controlling-shadow-ai/)

[![Top 5 AI Gateways for SSO and RBAC in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/top-5-ai-gateways-for-sso-and-rbac-in-2026-bifrost-isometric.png) AI gateways with SSO and RBAC let enterprises tie every model request and every configuration change to a corporate identity. This guide compares Bifrost, Kong AI Gateway, Azure API Management, Gravitee, and Cloudflare AI Gateway on identity, roles, provisioning, and audit.](https://www.getmaxim.ai/articles/top-5-ai-gateways-for-sso-and-rbac-in-2026/)

[![Semantic Caching: The Top 5 AI Gateways in 2026](https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w720/2026/10/semantic-caching-the-top-5-ai-gateways-in-2026-bifrost-isometric.png) Semantic caching serves a stored LLM response when a new prompt means the same thing as an earlier one. This guide compares Bifrost, Kong AI Gateway, Azure API Management, and Cloudflare AI Gateway on match modes, vector stores, thresholds, TTLs, and cache scoping.](https://www.getmaxim.ai/articles/semantic-caching-the-top-5-ai-gateways-in-2026/)

```json
{
    "@context": "https://schema.org",
    "@type": "Article",
    "publisher": {
        "@type": "Organization",
        "name": "Maxim Articles",
        "url": "https://www.getmaxim.ai/articles/",
        "logo": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w256h256/2025/08/thumbnail.png",
            "width": 60,
            "height": 60
        }
    },
    "author": {
        "@type": "Person",
        "name": "Kamya Shah",
        "image": {
            "@type": "ImageObject",
            "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/2025/09/WhatsApp-Image-2025-08-29-at-17.40.40-1.jpeg",
            "width": 1200,
            "height": 1600
        },
        "url": "https://www.getmaxim.ai/articles/author/kamya/",
        "sameAs": []
    },
    "headline": "What an LLM Gateway Actually Does: A Guide for AI Infrastructure Teams",
    "url": "https://www.getmaxim.ai/articles/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/",
    "datePublished": "2026-06-02T04:59:00.000Z",
    "dateModified": "2026-10-08T16:05:32.000Z",
    "image": {
        "@type": "ImageObject",
        "url": "https://storage.ghost.io/c/84/03/8403f2f6-141c-411a-8f55-a32d4291533e/content/images/size/w1200/2026/07/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastruct-bifrost-paper-waves.optimized.png",
        "width": 1200,
        "height": 630
    },
    "keywords": "AI Gateway",
    "description": "An LLM gateway runs a fixed pipeline on every model call, from virtual key authentication to routing, guardrails, and logging. This guide walks infrastructure teams through each stage, the deployment trade-offs, and how to evaluate one.",
    "mainEntityOfPage": "https://www.getmaxim.ai/articles/what-an-llm-gateway-actually-does-a-guide-for-ai-infrastructure-teams/"
}
```
