What Is an MCP Gateway? A Guide for Production AI Agents
An MCP gateway is a centralized layer between AI agents and the MCP servers they call, exposing every connected tool through one endpoint. This guide covers the four roles a gateway plays, how it differs from an MCP server and a proxy, the deny-by-default access model, and when to adopt one.
TL;DR
- An MCP gateway is a centralized layer between AI agents and the Model Context Protocol servers they use, exposing every connected tool through one endpoint with governance and a full audit trail.
- It solves the sprawl that appears at scale: ungoverned credentials, duplicated tool catalogs, and token costs that balloon as the number of MCP servers per agent grows from a handful to dozens.
- A production MCP gateway plays four roles at once: MCP client to external servers, MCP server to agents, governance layer, and observability layer.
- Bifrost acts as all four, with per-virtual-key tool filtering and a Code Mode execution path that cut input tokens by 92.8% in published benchmarks.
- Security is centralized rather than per-agent: OAuth 2.0, curated tool groups, and immutable audit logs at one control plane.
An MCP gateway is a centralized infrastructure layer that sits between AI agent clients and the Model Context Protocol (MCP) servers they connect to. Instead of every agent managing its own MCP server list, credentials, and tool catalogs, an MCP gateway exposes all connected tools through a single endpoint, enforces governance policies, and provides a complete audit trail of every tool execution. As AI agents move from prototypes into production and the number of MCP servers per agent grows from a handful to several dozen, an MCP gateway becomes the difference between a manageable, governable agent estate and an ungoverned sprawl of credentials, configurations, and token costs. Bifrost, the open-source AI gateway built by Maxim AI, was designed for exactly this control point, combining native MCP support with virtual-key governance, tool filtering, and a Code Mode execution path that has been shown to reduce input tokens by up to 92.8% at scale.
Why MCP Created the Need for a Gateway
The Model Context Protocol was introduced by Anthropic in November 2024 as an open standard for connecting AI assistants to the systems where data lives, including content repositories, business tools, and development environments, with the explicit goal of replacing fragmented one-off integrations with a single protocol. The architecture is straightforward: developers can either expose their data through MCP servers or build AI applications (MCP clients) that connect to these servers. Within twelve months, the Linux Foundation announced the formation of the Agentic AI Foundation, and Anthropic donated MCP to the foundation as a vendor-neutral standard.
That adoption curve created a new operational problem. Agents that worked fine with two or three MCP servers started showing strain once teams connected ten, twenty, or fifty. The pain points fall into three buckets:
- Connection sprawl: Every agent maintains its own list of MCP server endpoints, transports (STDIO, HTTP, SSE), and credentials.
- Token inflation: Default MCP clients inject every tool definition from every connected server into the model's context on every request.
- Governance blind spots: Without a central control plane, there is no consistent way to filter which tools a given consumer can call, set budgets, or audit what executed.
An MCP gateway is the infrastructure response to those three problems.

What an MCP Gateway Actually Does
A production-grade MCP gateway plays four distinct roles in the agent stack:
- MCP client: Connects to external MCP servers (filesystem, search, databases, custom APIs) and discovers their capabilities at startup.
- MCP server: Exposes the entire connected tool ecosystem through a single gateway URL that any MCP-aware client (Claude Desktop, Claude Code, Cursor, Codex CLI) can point at.
- Policy enforcement point: Applies access controls, rate limits, budget caps, and tool allow-lists before any tool executes.
- Audit and observability layer: Logs every tool suggestion, approval, and execution with full metadata for compliance and debugging.
| Role | What it faces | What it controls |
|---|---|---|
| MCP client | Upstream MCP servers | Connection, transport (STDIO, HTTP, SSE), credential storage, capability discovery |
| MCP server | Downstream agents and IDEs | One endpoint, one tool catalog, one credential per consumer |
| Policy enforcement point | Both directions | Tool allow-lists, rate limits, budget caps, guardrail evaluation |
| Observability layer | Both directions | Request logs, token and cost attribution, latency per tool call |
Bifrost is built to act as all four simultaneously. Bifrost's MCP gateway auto-discovers tools from any MCP-compatible server, injects them into chat completion requests, and returns tool suggestions back to the calling application, all while keeping execution stateless by default so applications retain full control over which calls actually run.

MCP Gateway vs MCP Server vs MCP Proxy
The three terms describe different layers. An MCP server exposes one set of tools over the protocol. An MCP proxy forwards traffic to a server, usually to bridge a transport or add a credential. An MCP gateway terminates the protocol on both sides, so it is a client to many servers and a server to many agents, and it enforces policy in the middle.
| MCP server | MCP proxy | MCP gateway | |
|---|---|---|---|
| Faces agents | Yes, directly | Yes, as a pass-through | Yes, as a single endpoint |
| Faces other servers | No | One, usually | Many, concurrently |
| Aggregates tool catalogs | No | No | Yes |
| Enforces access policy | Per server, if at all | Rarely | Centrally, per consumer |
| Records executions | Locally, if at all | Transport-level only | Every call, with attribution |
| Typical deployment | One per tool or data source | Sidecar or bridge | One per environment |
The distinction matters operationally because a proxy scales with the number of servers while a gateway scales with the number of environments. Teams evaluating the boundary in more detail can work through how an MCP gateway, proxy, and server differ in practice, and MCP proxy architecture and its use cases covers the cases where a proxy is genuinely the right layer.
The Tool Sprawl Problem MCP Gateways Solve
The cost of an ungoverned MCP estate is not theoretical. Independent security analysis has documented more than 10,000 active public MCP servers, alongside seven concrete enterprise security risks tied to that growth: sensitive data exfiltration, unauthorized agent actions, overprivileged access, supply chain exposure, missing audit trails, privilege escalation, and shadow AI sprawl.
Cloudflare, which operates MCP at company-wide scale, has been explicit about why locally hosted MCP servers are a liability: local MCP server deployments may rely on unvetted software sources and versions, which increases the risk of tool injection attacks, and they prevent IT and security administrators from administrating these servers, leaving it up to individual employees and developers to choose which MCP servers they want to run. The conclusion is straightforward: when MCP moves from a single developer's laptop to a production fleet, a gateway is the only place where consistent controls can be enforced.
Bifrost addresses this with tool filtering per virtual key, where the default behavior is deny-by-default. A virtual key with no MCP configuration sees no MCP tools at all, so access is an explicit allow-list rather than an implicit grant. Administrators configure which MCP clients and which specific tools each consumer can reach through the governance layer.
The failure mode this prevents has a name: shadow MCP, where ungoverned servers run outside any inventory, and the broader policy model is covered in what MCP governance means in practice.
How an MCP Gateway Controls Cost at Scale
The cost problem with classic MCP is structural, not incidental. Most MCP clients load all tool definitions upfront directly into context, exposing them to the model using a direct tool-calling syntax, which means tool descriptions occupy more context window space, increasing response time and costs. Anthropic's own engineering team has documented the scale of this problem and proposed code execution as the response, noting that MCP provides a foundational protocol for agents to connect to many tools and systems, but once too many servers are connected, tool definitions and results can consume excessive tokens, reducing agent efficiency.
Bifrost's response is Code Mode, an execution path where the model writes a short Python (Starlark) script that orchestrates multiple tool calls inside a sandbox, rather than calling each tool individually through the model's tool-calling loop. Instead of injecting hundreds of tool schemas on every turn, Code Mode exposes only four meta-tools:
- listToolFiles: enumerates the MCP servers and tools that are reachable
- readToolFile: returns Python function signatures for a chosen server or tool
- getToolDocs: fetches detailed documentation for a specific tool on demand
- executeToolCode: runs an orchestration script inside a sandboxed Starlark interpreter
Code Mode was benchmarked against classic MCP across three rounds of increasing MCP footprint, with the same query set run twice per round:
| Round | MCP footprint | Pass rate, classic | Pass rate, Code Mode | Input tokens, classic | Input tokens, Code Mode | Change |
|---|---|---|---|---|---|---|
| 3 | 508 tools / 16 servers | 100% | 100% | 75.1M | 5.4M | -92.8% |
t roughly 500 tools, average input tokens per query fell from 1.15M to 83K, a reduction of about 14x, with estimated cost falling 92.2% and execution running around 40% faster. Pass rate held at 100% in both configurations, which matters more than the token number: a cost optimization that degrades task completion is not an optimization. The savings curve compounds as the MCP footprint grows. Full methodology and results are published in the Bifrost performance benchmarks, and how code execution cuts agent token costs at scale as its own topic. Teams comparing gateways specifically on cost behavior can start from the best MCP gateways for production AI systems.

MCP Gateway Security: Why a Central Control Plane Matters
MCP security is not a niche concern. In April 2025, security researchers released an analysis that concluded there are multiple outstanding security issues with MCP, including prompt injection, tool permissions that allow for combining tools to exfiltrate data, and lookalike tools that can silently replace trusted ones. The official MCP specification itself acknowledges the trust model that implementations must build on top, noting that tools represent arbitrary code execution and must be treated with appropriate caution, descriptions of tool behavior such as annotations should be considered untrusted unless obtained from a trusted server, and hosts must obtain explicit user consent before invoking any tool.
A gateway is where those abstract principles become enforceable controls. Bifrost ships several primitives explicitly designed for this:
- Stateless execution by default: Tool calls returned by the LLM are suggestions, not actions. Execution requires an explicit, separately authorized API call.
- OAuth 2.0 with PKCE: Secure delegated authentication with automatic token refresh for upstream services, scoped per end-user.
- Tool groups and virtual keys: Named collections of tools attached to virtual keys, teams, or customers, resolved at request time with no implicit access.
- Request logs: Every tool call recorded with inputs, outputs, latency, token usage, and cost, filterable by tool name so a specific function can be traced across an estate.
- Audit logs: Administrative activity recorded as signed events, capturing the initiator, the action, the target resource, and the outcome, with configurable retention and archival to S3 or GCS.
- Enterprise guardrails: Three Bifrost-managed checks (Prompt Guardrails, Custom Regex, Secrets Detection) plus eleven external providers including Presidio, Azure AI Language PII, AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, CrowdStrike AIDR, Gray Swan, and Patronus AI. Rules are written in CEL and evaluate MCP tool arguments and results, not only LLM prompts and responses.
Authentication deserves particular attention because it is where per-agent credential sprawl actually lives. Bifrost supports six auth types for upstream MCP servers, and the right one depends on whether the credential belongs to the organization or to the individual user:
| Auth type | Who authenticates | When to use |
|---|---|---|
none | Nobody | Local or unauthenticated servers |
headers | Administrator, once | Shared API keys and bearer tokens |
oauth | Administrator, once | A third-party service the whole team shares |
per_user_headers | Each end-user, lazily | Per-user API keys or signed tokens |
per_user_oauth | Each end-user, lazily | Per-user services such as Notion, GitHub, or Sentry |
token_exchange | Each caller, per request | Caller identity forwarded with no persisted per-user credential |
Per-user auth applies to HTTP and SSE connections; STDIO connections inherit their environment from the spawned subprocess and have no per-call auth model. The full decision path is covered in MCP authentication patterns for OAuth, API keys, and token management, and the controls above are enumerated as a checklist in MCP security best practices for enterprise deployments.
For regulated industries, these controls map directly onto compliance requirements. Teams in financial services, healthcare, and the public sector can review Bifrost's vertical-specific patterns in the healthcare AI infrastructure and financial services pages.
What Is MCP Gateway Architecture, and Where Does It Sit in the Agent Stack?
MCP gateway architecture places the gateway as a horizontal layer between agent clients above and MCP servers below, terminating the protocol on both sides. It does not replace the LLM provider, the agent framework, or any individual MCP server. It consolidates the connection, credential, and policy decisions that would otherwise be duplicated inside every agent.
An MCP gateway is a horizontal layer, not a replacement for any specific component. It sits between:
- Upstream: AI agent clients (Claude Code, Cursor, Codex CLI, custom applications) and the LLM providers they use.
- Downstream: MCP servers exposing tools, data sources, and APIs (filesystem, databases, internal services, third-party integrations).
The same gateway typically also handles LLM routing, fallbacks, and observability for the model side of the same workloads. An MCP gateway becomes essential, functioning as a production LLM gateway that centralizes tool discovery, routing, governance, and execution so workflows remain predictable and debuggable. Bifrost is built on this principle: a single Go-based control plane adds only 11 microseconds of overhead per request at 5,000 RPS, while consolidating provider routing, MCP tool execution, and governance into one deployment.
Wiring existing MCP clients into Bifrost requires a single configuration change: point the client's MCP endpoint at the gateway URL. Every connected MCP server becomes reachable through that single URL, scoped by whichever virtual key is in use. The Bifrost Claude Code integration covers this setup for one of the most common agent surfaces.

How to Connect MCP Clients to the Gateway
Connecting a client is a configuration change, not a code change. An MCP-aware client points its server endpoint at the gateway URL and authenticates with a virtual key; the gateway then presents every upstream server that key is allowed to reach as though it were a single server.
Bifrost exposes two endpoints for this. POST /mcp carries JSON-RPC 2.0 messages for tool discovery and execution, and GET /mcp opens a Server-Sent Events stream for clients that hold a persistent connection. Every request to /mcp is scoped by the credential it carries, so two teams pointing at the same URL see different tool catalogs.
Clients can authenticate with a header credential or through a browser-based OAuth flow, controlled by the mcp_server_auth_mode setting. Coding agents are the most common first surface: adding and governing MCP servers in Claude Code walks through that setup, and the Bifrost Claude Code resource page covers the configuration reference. Where the upstream servers are hosted SaaS products rather than local processes, connecting remote MCP servers through one gateway covers the transport and credential differences.
Discovery across a large estate is its own problem once the server count grows. An MCP registry indexes which servers exist and who may use them, and it complements rather than replaces the gateway: the registry answers what is available, the gateway decides what is reachable.
MCP Gateway Observability and Audit Trails
An MCP gateway is the only point in an agent stack where every tool call is visible, because it is the only component every call passes through. Observability at the gateway answers three questions that per-agent logging cannot: which tools ran, what they cost, and who authorized them.
Bifrost separates these into two record types, and the distinction matters for compliance reviews:
| Request logs | Audit logs | |
|---|---|---|
| Records | LLM calls and MCP tool executions | Administrative activity |
| Captured fields | Inputs, outputs, tokens, cost, latency, tool names | Initiator, action, target resource, outcome |
| Answers | What did the agent do | Who changed the configuration |
| Filtering | By latency, token range, tool call name, virtual key | By initiator, target, IP, action, outcome, date |
| Retention | Configurable, with content logging optionally disabled | Configurable, with archival to S3 or GCS |
Request logging runs asynchronously and does not add latency to the request path. Content logging can be disabled independently, so usage metadata such as cost, token count, and latency is still recorded when the payloads themselves cannot be retained. Auditing every AI tool call at the gateway covers the review workflow this enables, and how MCP tools are discovered, invoked, and access-controlled explains what each field represents.
When to Adopt an MCP Gateway
Most teams adopt an MCP gateway when the number of MCP servers passes roughly three per agent, when a second team starts consuming the same tools, or when an agent begins touching regulated data. Before those thresholds, per-agent configuration is manageable. After them, it diverges faster than it can be reconciled by hand.
The decision point is usually a function of three signals:
- Server count: Beyond three MCP servers, the token-cost and tool-selection problems become noticeable. Beyond ten, they become unavoidable.
- Team count: As soon as more than one team consumes MCP tools, divergence in configuration, credentials, and policies starts. A gateway is the only durable answer.
- Production readiness: Any agent moving from prototype to a customer-facing or revenue-affecting workflow needs auditable tool execution. Local MCP servers cannot provide that.
| Signal | Below the threshold | Above the threshold |
|---|---|---|
| MCP servers per agent | 1-3: per-agent config is fine | 10+: token cost and tool selection degrade |
| Teams consuming tools | One: conventions hold | Two or more: credentials and policy diverge |
| Data sensitivity | Internal or synthetic | Regulated or customer-facing: execution must be auditable |
Teams comparing deployment options at this stage can review the best open-source MCP gateways alongside commercial platforms.
MCP itself is now governed under the Linux Foundation, with broad cross-vendor adoption from OpenAI, Google, Microsoft, AWS, and Cloudflare signaling that the protocol is here to stay; the operational question is no longer whether MCP, but how to run it safely.
Frequently Asked Questions
What is the difference between an MCP gateway and an AI gateway?
An AI gateway governs traffic to LLM providers: routing, failover, budgets, and caching for model calls. An MCP gateway governs traffic to tools: discovery, access control, and audit for the Model Context Protocol servers agents call. Bifrost is both, so model access and tool access are governed through one control plane.
Why do production AI agents need an MCP gateway?
Without one, each agent holds its own MCP server list and credentials, and no team has a complete inventory of what those agents can call. That becomes ungovernable as server counts grow. An MCP gateway centralizes discovery, enforces which tools each consumer can reach, and logs every execution, turning a sprawl of connections into a governable estate.
How does an MCP gateway reduce token costs?
Classic MCP loads every tool definition into the model's context on every call, so cost scales with the size of the tool catalog. Bifrost's Code Mode instead has the model write code that enumerates and calls tools on demand, which reduced input tokens by 92.8% across 508 tools in controlled benchmarks, with the savings compounding as the MCP footprint grows.
What is MCP Code Mode?
Code Mode is an execution path where the agent writes code to discover and invoke MCP tools rather than having every tool definition injected into its prompt. It enumerates available servers and tools, calls only what a task needs, and returns results, which keeps context small and cost low. Bifrost implements Code Mode natively in the gateway.
How does an MCP gateway improve security?
It replaces per-agent credentials and ad-hoc tool access with a central control plane. OAuth 2.0 handles authentication, tool groups scope which tools each virtual key can reach, stateless execution avoids persisting tool state, and immutable audit logs record every call for compliance review.
When should a team adopt an MCP gateway?
The usual trigger is scale on three axes: more than a few MCP servers per agent, multiple teams sharing tools, or regulated data flowing through tool calls. Past roughly three servers the token-cost and governance problems stop being theoretical, and a gateway becomes cheaper to operate than the sprawl it replaces.
Getting Started with an MCP Gateway
An MCP gateway turns the Model Context Protocol from a useful integration standard into infrastructure that scales. The adoption path is usually incremental rather than a migration: point one agent surface at the gateway, scope it with a virtual key, confirm the tool catalog and the request logs look right, then move the remaining clients over one at a time. Nothing about the upstream MCP servers changes, which is what makes the first step cheap to reverse.
Bifrost is a single control plane for MCP tool discovery, governance, audit, and cost optimization. It unifies 25+ providers and 10,000+ models through one OpenAI-compatible API, adds 11 microseconds of overhead per request at 5,000 RPS, supports Code Mode natively to cut token usage by up to 92.8% as the tool surface grows, and ships an open-source core under Apache 2.0. The same deployment governs model access and tool access, so an agent estate has one place to enforce policy rather than two.
To see how an MCP gateway fits into a production agent stack, book a demo with the Bifrost team.