The Best Open Source AI Gateway in 2026
An open source AI gateway is a self-hostable service under an open-source license that routes, governs, and observes LLM traffic through one API. This guide compares Bifrost, LiteLLM, Kong AI Gateway, Envoy AI Gateway, agentgateway, and Apache APISIX on license scope and benchmarks.
TL;DR
- Open source AI gateways have become essential infrastructure for production LLM applications. The strongest 2026 options publish their code under permissive licenses such as Apache 2.0 or MIT.
- Bifrost ranks first: written in Go and licensed under Apache 2.0, it adds 11 µs of overhead per request at 5,000 RPS, with built-in MCP support, semantic caching, governance, and zero-config startup.
- License scope matters as much as the license name: LiteLLM and Kong keep features such as admin SSO, audit logs, or semantic caching in a commercial tier (open core).
- Of the six open source AI gateways compared here, only Bifrost and LiteLLM state a performance figure in their README or docs, and Bifrost publishes its benchmarking tool so the result can be reproduced.
An open source AI gateway is a self-hostable service, published under an open-source license, that sits between applications and LLM providers to handle routing, failover, caching, access control, and observability through one API. Bifrost, the open source AI gateway released on GitHub under the Apache 2.0 license, is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. This guide ranks six open source AI gateways on the criteria specific to open-source software: license and open core scope, source transparency, community activity, extensibility, and published benchmarks.
Why AI Gateways Matter More Than Ever
As AI systems move from proof-of-concept to production, teams quickly realize that calling LLM providers directly from application code creates a fragile, expensive, and ungovernable architecture. Each provider has its own API format, authentication scheme, rate limits, and pricing model. Multiply that across OpenAI, Anthropic, AWS Bedrock, Google Vertex, Mistral, and others, and the integration burden compounds as usage scales.
An AI gateway sits between your application and these providers, acting as a unified control plane for routing, failover, caching, cost management, and observability (the AI gateway architecture guide covers each layer). In 2026, with enterprises rapidly shifting from pilot programs to full production deployments, the gateway layer is no longer optional middleware. It is core infrastructure: Deloitte's 2026 State of AI in the Enterprise report expects the number of companies with 40% or more of their AI projects in production to double within six months.

Figure 1: Because the gateway runs from source in your own infrastructure, every policy it applies can be read, audited, and changed.
The practical question is which open source gateway to trust with production traffic, and an open-source license alone does not answer it.
How to Choose an Open Source LLM Gateway
Choose an open source LLM gateway by checking five things beyond features: the license on the code you will actually run, which capabilities sit in a commercial tier, who stewards the project, how it is extended, and whether its performance claims are published and reproducible.
| Criterion | What to check | Why it matters |
|---|---|---|
| License | An OSI-approved license (Apache 2.0, MIT) on every directory you deploy | Decides whether you can self-host, modify, and redistribute without a subscription |
| Open core scope | Separate enterprise/ directories, license keys, or "Enterprise" labels in the docs |
SSO, audit logs, or caching can be gated even when the repository is public |
| Steward and community | Company, Linux Foundation, or Apache Software Foundation governance; stars, forks, release cadence | Signals longevity and how fast fixes land |
| Extensibility | Plugin model and supported languages | Custom auth, redaction, or routing without forking the gateway |
| Published benchmarks | Overhead figure, test hardware, and whether the test tooling is open source | Lets you verify a performance claim on your own infrastructure |
The top open source LLM gateways compared hub scores the same category on routing and deployment depth; this guide scores it on open-source criteria. Self-hosting teams can cross-check the open source LLM gateways for self-hosted deployments.
The Open Source Landscape in 2026
Six open source AI gateways cover most production shortlists in 2026: Bifrost, LiteLLM, Kong AI Gateway, Envoy AI Gateway (now Agent Router), agentgateway, and Apache APISIX. All six publish source code under Apache 2.0 or MIT, but they differ in language, steward, open core scope, extension model, and whether they state a performance figure.
Several open source AI gateways have gained traction over the past year. Here is how the most prominent options stack up. Every cell comes from the project's own repository, license file, or docs as read in September 2026.
| Gateway | License | Language | Steward | GitHub stars | Extension model | Performance figure in README or docs |
|---|---|---|---|---|---|---|
| Bifrost | Apache 2.0 | Go | Company-backed | ~7.5k | Native Go plugins on request-lifecycle hooks | 11 µs overhead at 5,000 RPS |
| LiteLLM | MIT, except the enterprise/ directory (commercial license) |
Python | BerriAI | ~53.8k | Python callbacks | 8 ms P95 latency at 1,000 RPS |
| Kong AI Gateway | Apache 2.0 (Kong Gateway); some AI plugins Enterprise-only | Lua | Kong Inc. | ~43.9k (whole API gateway) | Lua, Go, and JavaScript plugins | None stated in README |
| Envoy AI Gateway (Agent Router) | Apache 2.0 | Go | Agentic AI Foundation | ~1.8k | Kubernetes CRDs on Envoy Gateway | None stated in README |
| agentgateway | Apache 2.0 | Rust | Linux Foundation | ~4.6k | CEL policy engine | None stated in README |
| Apache APISIX | Apache 2.0 | Lua | Apache Software Foundation | ~17.1k (whole API gateway) | Lua; Java, Go, Python, Node.js via RPC; Wasm | None stated in README |
Star counts are rounded, and Kong and APISIX count stars for the whole API gateway. The two published figures measure different things on different hardware (gateway overhead versus end-to-end P95 latency), so they are not a head-to-head result; the benchmarks section below explains how to produce one. For a five-gateway breakdown focused on routing and deployment depth, see how the leading open-source LLM gateways compare.
Why Bifrost Stands Out
Bifrost ranks first because it scores well on every open-source criterion at once: an Apache 2.0 license on the gateway core, public source in Go, a native plugin system, and a published overhead figure (11 µs at 5,000 RPS) backed by open source benchmarking tools.
Bifrost is an open source AI gateway, written in Go and built by Maxim AI. Bifrost is designed for teams where latency, throughput, reliability, and governance are hard requirements rather than aspirations.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

Figure 2: Built-in features and custom Go plugins attach to the same hooks, so extending the open source gateway never requires a fork.
Performance That Holds Up Under Load
Bifrost adds just 11 microseconds of overhead per request at 5,000 RPS with a 100% success rate. The figure comes from sustained tests on a t3.xlarge instance (4 vCPUs, 16 GB) against mocked OpenAI calls. Gateway overhead matters in production: when you have hundreds of developers making thousands of requests per day, or multi-step agentic workflows chaining several model calls in sequence, microsecond-level gateway overhead directly affects user experience and system throughput.
Go's native concurrency model, optimized connection pooling, and minimal processing footprint make this possible. Bifrost was engineered for high-throughput, long-running production workloads from day one.
Zero-Config Startup, Full Production Depth
Getting started takes one command:
npx -y @maximhq/bifrost
No configuration files. No environment setup. Bifrost launches with a Web UI for visual configuration, real-time monitoring, and analytics by default, and the same gateway runs as a container with docker run -p 8080:8080 maximhq/bifrost. For teams that prefer infrastructure-as-code, file-based and API-driven configuration are fully supported.
Unified Interface Across 25+ Providers
Bifrost provides a single OpenAI-compatible API that routes to OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Cohere, Mistral, Groq, Ollama, and more: 25+ providers and 10,000+ models in total. Switching providers or adding fallbacks requires no code changes. Replace your existing SDK base URL with Bifrost's endpoint and Bifrost works as a drop-in replacement, and automatic fallback routing becomes a configuration change rather than a code change.
Built-In MCP Gateway
As agentic AI systems become more prevalent, models need to interact with external tools: filesystems, web search, databases, and third-party APIs. Bifrost includes a native MCP (Model Context Protocol) gateway that centralizes tool connections, governance, security, and authentication. Instead of managing MCP integrations at the application layer, teams can enforce policies at the infrastructure level. MCP Code Mode cut input tokens by 58.2% at 96 tools and by 92.8% at 508 tools in published benchmarks, which is why Bifrost also appears among the best open source MCP gateways.
Semantic Caching
Bifrost's semantic caching identifies semantically similar requests and serves cached responses, reducing both latency and cost without sacrificing response quality. The same plugin also supports direct hash matching for identical requests, which needs no embedding provider. For workloads with repeating or near-duplicate queries, this can drive meaningful savings at scale.
Governance in the Open Source Core
Production AI at scale requires more than routing. It requires cost controls, access management, and audit trails. Bifrost delivers hierarchical budget management with virtual keys, rate limits per virtual key and provider, and team- and customer-level spend caps, and native Prometheus metrics plus OpenTelemetry distributed tracing are part of the same open source gateway.
Virtual keys, budgets and rate limits, routing, and MCP tool filtering all ship under Apache 2.0, and the Bifrost governance model needs no license key. Bifrost Enterprise is a strict superset with the same config.json schema that adds OIDC single sign-on, RBAC, clustering, guardrails, audit logs, and secret management with HashiCorp Vault, AWS Secrets Manager, or GCP Secret Manager. No vendor lock-in.
Extensible with Native Go Plugins
Bifrost is extended through native Go plugins that load as shared objects in a dynamically linked build and attach to hooks across the request lifecycle, as Figure 2 shows: HTTP transport hooks, a once-per-request routing hook that can choose the provider, model, and fallbacks, and pre- and post-LLM hooks that run on every provider attempt. Built-in features such as governance and semantic caching are configured as plugins on the same plugin architecture, so custom logic runs in-process with no IPC hop. WASM plugins are deprecated, and webhook-based plugins are planned as the cross-language path.
Other Open Source AI Gateways Worth Evaluating
Five other open-source projects are worth a place on a 2026 shortlist: LiteLLM for provider breadth in Python, Kong AI Gateway and Apache APISIX for teams already running those API gateways, and Envoy AI Gateway (now Agent Router) and agentgateway for Kubernetes-native and agent-protocol traffic.
2. LiteLLM
LiteLLM remains the most widely adopted Python-based gateway, with support for 100+ LLM providers. Its strength is breadth: if you need to connect to a niche provider, LiteLLM likely supports it. The repository's license file makes everything outside the enterprise/ directory MIT-licensed, while that directory carries a separate BerriAI license that requires a subscription for production use. LiteLLM's docs list SSO for the admin UI, audit logs, JWT authentication, and secret manager integrations as Enterprise features, and its README states 8 ms P95 latency at 1,000 RPS.
Best for: Python-heavy teams that want the broadest provider catalog and accept an open core license split. The Bifrost vs LiteLLM enterprise comparison and the LiteLLM alternatives page cover the trade-offs in depth.
3. Kong AI Gateway
Kong AI Gateway is the AI feature set of Kong Gateway, an Apache 2.0 API gateway written in Lua. The open source repository ships the AI Proxy, AI Prompt Guard, AI Prompt Template, AI Prompt Decorator, and AI request and response transformer plugins, while AI Semantic Cache is documented as available only in AI Gateway Enterprise. Custom plugins can be written in Lua, Go, or JavaScript.
Best for: teams already running Kong for REST traffic that want to add LLM routing to the same plane. The Kong AI Gateway alternatives guide compares it with purpose-built AI gateways.
4. Envoy AI Gateway (Now Agent Router)
Envoy AI Gateway has been renamed Agent Router and is now an Agentic AI Foundation project, with the same code, maintainers, and Apache 2.0 license. Written in Go on top of Envoy and Envoy Gateway, it uses a two-tier pattern: a first-tier gateway for authentication, top-level routing, and global rate limiting, and a second tier for self-hosted model serving. It runs standalone through the aigw CLI or on Kubernetes, and reached v1.0.0 in June 2026.
Best for: platform teams standardized on Envoy and the Kubernetes Gateway API. See the Envoy AI Gateway alternatives for teams that do not run Envoy.
5. agentgateway
agentgateway is an Apache 2.0 gateway written in Rust and governed as a Linux Foundation project, built around the MCP and A2A agent protocols. Its README lists an LLM gateway with budget and spend controls, load balancing, and failover, an MCP gateway with tool federation, and RBAC through a CEL policy engine. It runs as a standalone binary or on Kubernetes with the Gateway API.
Best for: teams whose main traffic is agent-to-tool and agent-to-agent rather than application-to-LLM.
6. Apache APISIX
Apache APISIX is an Apache Software Foundation API gateway, written in Lua and licensed under Apache 2.0, that acts as an AI gateway through plugins such as ai-proxy, ai-proxy-multi, and ai-rate-limiting. Its README lists LLM load balancing, retries, fallbacks, token-based rate limiting, and an mcp-bridge plugin, with external plugins in Java, Go, Python, and Node.js plus Wasm.
Best for: teams that already operate APISIX and want foundation-governed, vendor-neutral stewardship.
A Note on Cloudflare AI Gateway
Cloudflare AI Gateway runs on Cloudflare's global network with caching, rate limiting, spend limits, and analytics. It is available on all Cloudflare plans and convenient for teams already on the Cloudflare stack, but it is a managed service with no self-hosted option and no published source, so it does not meet the open-source criteria in this guide. The Cloudflare AI Gateway alternatives article covers it separately.
Each of these tools solves a real problem. But when the criteria shift from "easy to set up" to "reliable under sustained production load," the architectural choices behind the gateway start to matter significantly.
Open Source vs Open Core: What the License Covers
Open source means the code you deploy is under an OSI-approved license with no feature held back; open core means a permissive core plus features that require a commercial license or subscription. Several open source AI gateways are open core, so the practical check is which capabilities you need and which side of the line each one sits on.

Figure 3: Reading the license file is the start of the check; the answer is in the directories and tier labels behind it.
| Gateway | In the permissive core | In a commercial tier |
|---|---|---|
| Bifrost | Routing, fallbacks, load balancing, virtual keys, budgets, rate limits, semantic caching, MCP gateway, Prometheus, OpenTelemetry, Go plugins | Clustering, adaptive load balancing, guardrails, OIDC SSO, RBAC, audit logs, secret management, in-VPC deployment |
| LiteLLM | Proxy, virtual keys, spend tracking, load balancing, guardrails (README) | Admin UI SSO, audit logs, JWT auth, secret managers, IP access lists (Enterprise license) |
| Kong AI Gateway | AI Proxy, prompt guard, template, and decorator, request and response transformers | AI Semantic Cache and other AI Gateway Enterprise plugins |
| Envoy AI Gateway, agentgateway, Apache APISIX | Full repository under Apache 2.0 | No commercial tier stated in the repository |
The code in all six repositories is covered by the Apache License 2.0 or MIT, both of which permit commercial use, modification, and self-hosting. The difference is what you pay for later: with Bifrost, everything in the first column is free under Apache 2.0, and Enterprise adds scale and compliance controls rather than gating basic governance. The Bifrost LLM gateway buyer's guide lists the questions to ask any vendor about license scope.
Published Benchmarks and How to Reproduce Them
A published benchmark is only useful if you can rerun it. Bifrost states 11 µs of gateway overhead at 5,000 RPS and publishes the tooling that produced it, so any team can measure Bifrost, or Bifrost next to another gateway, on its own hardware with a mock provider that removes model latency from the result.

Figure 4: A mock provider removes model latency from the measurement, which is how a gateway overhead figure such as 11 µs becomes reproducible.
The Bifrost benchmarking tool is open source and ships two companions: a mocker that simulates an LLM provider with configurable latency, failures, and rate limits, and a hitter that generates multi-model and streaming load. A run takes a fixed rate or concurrency level and can target other supported gateways as well as Bifrost. The published results cover two EC2 instance sizes: overhead fell from 59 µs on a t3.medium to 11 µs on a t3.xlarge at 5,000 RPS, both with a 100% success rate.
For throughput-focused shortlists, the open source AI gateways for high-throughput workloads comparison and the three-way OpenRouter, LiteLLM, and Bifrost comparison apply the same numbers to specific traffic patterns.
When to Choose Bifrost
The Bifrost AI gateway is the right choice if you are running high-traffic, customer-facing AI systems where latency and reliability matter. It also fits if compliance requirements such as GDPR, HIPAA, or SOC 2 Type II call for self-hosted deployment. If your engineering team is scaling beyond a handful of developers and needs per-team budget controls and access management, virtual keys and hierarchical spend controls cover that in the open source core. It also fits if you want a gateway that grows with you from prototype to production without requiring a migration.
If your primary need is broad provider experimentation in a Python-heavy environment with low traffic, LiteLLM remains a reasonable starting point. If you are deeply invested in Cloudflare's ecosystem, their AI Gateway is a natural extension. But for production-grade performance, governance, and reliability under a permissive license, Bifrost ranks first among the open source AI gateways in this comparison.
Frequently Asked Questions
Which open-source AI gateway is the best?
Bifrost is the best open source AI gateway for production traffic in 2026 on the criteria in this guide. It is Apache 2.0 licensed, written in Go, adds 11 µs of overhead per request at 5,000 RPS, and includes virtual keys, budgets, semantic caching, and an MCP gateway in the open source core. LiteLLM is the common alternative for Python-first teams that prioritize provider breadth.
What is an AI gateway?
An AI gateway is a service that sits between applications and LLM providers and exposes one API for routing, failover, caching, access control, cost tracking, and observability. Applications send every model request to the gateway, which applies policy and forwards the call to the right provider. Bifrost is an AI gateway that exposes an OpenAI-compatible API across 25+ providers and 10,000+ models.
Is Bifrost open source?
Yes. Bifrost is released under the Apache 2.0 license, and the gateway core, including routing, fallbacks, virtual keys, budgets, rate limits, semantic caching, the MCP gateway, and Go plugins, is free to use, modify, and self-host. Bifrost Enterprise is a separate commercial superset that adds clustering, guardrails, OIDC SSO, RBAC, audit logs, and secret management with the same configuration schema.
Is there a self-hosted AI gateway available?
Yes. Every open source AI gateway in this guide can be self-hosted: Bifrost runs with one npx or Docker command and on Kubernetes via Helm, and LiteLLM, Kong, Envoy AI Gateway, agentgateway, and Apache APISIX all publish self-hosted builds. Cloudflare AI Gateway is the exception in this article because it is a managed service with no self-hosted option.
Is there a free open source AI gateway?
Yes. Bifrost, Envoy AI Gateway, agentgateway, and Apache APISIX are free under Apache 2.0, and LiteLLM is free under MIT outside its enterprise/ directory. Bifrost's open source core includes governance and semantic caching, so budgets and caching do not require a paid tier.
Is Kong open source?
Kong Gateway is open source under the Apache 2.0 license, and its repository includes AI plugins such as AI Proxy, AI Prompt Guard, and AI request and response transformers. Some Kong AI Gateway features, including AI Semantic Cache, are documented as available only in AI Gateway Enterprise, so Kong is best described as open core for AI workloads.
Get Started
Bifrost is open source, free to self-host under Apache 2.0, and starts with a single npx or Docker command. Run it locally, add provider keys in the Web UI, point an existing OpenAI or Anthropic SDK at the gateway, and measure overhead on your own hardware with the open source benchmarking tool.
- Documentation: docs.getbifrost.ai
- Website: getmaxim.ai/bifrost
To see how an open source AI gateway fits your production stack, from license scope to governance and benchmarks on your own hardware, book a demo with the Bifrost team.