TensorZero Alternatives: 5 LLM Gateways Compared in 2026
TensorZero is an open-source LLMOps stack built around a Rust LLM gateway, and its site now says it is no longer maintained. This guide compares five replacements, including Bifrost, LiteLLM, Kong AI Gateway, and Cloudflare AI Gateway, on routing, governance, MCP, and deployment.
TL;DR
- The TensorZero homepage now states that TensorZero remains available on GitHub but is no longer maintained, so production teams running its gateway need a supported replacement.
- TensorZero combined an LLM gateway with an optimization stack (evaluations, A/B tests, fine-tuning); most replacements are gateways first and treat optimization as a separate concern.
- Bifrost, an open-source AI gateway written in Go, adds 11 microseconds of overhead per request at 5,000 RPS and unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API.
- Teams that need virtual keys, hierarchical budgets, MCP tool governance, and in-VPC or air-gapped deployment should evaluate Bifrost first.
- LiteLLM, Kong AI Gateway, Agent Router (formerly Envoy AI Gateway), and Cloudflare AI Gateway each fit narrower cases: a Python-ecosystem proxy, an existing Kong or Envoy estate, or a managed edge service.
TensorZero is an open-source LLMOps stack that pairs a Rust LLM gateway with observability, evaluations, optimization, and experimentation, and its homepage now says the project is no longer maintained. Teams that adopted TensorZero mainly for its gateway need a supported LLM gateway that handles routing, governance, MCP, and enterprise deployment rather than another all-in-one optimization stack. Bifrost, the open-source AI gateway built in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. This guide compares Bifrost with four other TensorZero alternatives and shows what a migration involves.
What Is TensorZero and Why Teams Look for Alternatives
TensorZero is an open-source stack for LLM applications that unifies a gateway, observability, optimization, evaluations, and experimentation in one self-hosted deployment. Its homepage now states that TensorZero remains available on GitHub but is no longer maintained, which leaves production users without security fixes, provider updates, or a support path.
Before that change, the TensorZero README described five components:
- Gateway: a unified API for LLM providers, written in Rust, with a published figure of under 1 ms p99 latency overhead at 10,000+ QPS
- Observability: inferences and feedback stored in Postgres or ClickHouse and surfaced in the TensorZero UI
- Evaluations: inference-level and workflow-level evaluations using heuristics and LLM judges
- Optimization: supervised fine-tuning, automated prompt engineering, and inference-time strategies such as dynamic in-context learning
- Experimentation: adaptive A/B tests across prompts, models, and inference strategies
The design centered on a learning loop: store every inference, attach feedback, and use that data to improve prompts and models. Gateway-level controls existed but were narrower. TensorZero API keys required Postgres, custom rate limits ran on Valkey or Postgres, and the docs noted that UI authentication was "coming soon."

Figure 1: TensorZero centers on a feedback loop over stored inferences, while a gateway-first alternative centers on access, cost, and tool policy at request time.
Teams look for TensorZero alternatives for three reasons.
First, an unmaintained gateway in the request path carries operational risk as provider APIs change. Second, many teams adopted TensorZero for its gateway and used little of the optimization stack. Third, platform teams increasingly need controls TensorZero did not center on: per-team budgets, identity-based access, MCP tool governance, and private-network deployment. For background on the category, see this guide on what an LLM gateway does in enterprise AI.
Key Criteria for Choosing an LLM Gateway
An LLM gateway replacing TensorZero should be judged on what sits in the request path: overhead under load, provider coverage, routing and failover, access and cost governance, MCP support, and deployment options. Optimization features matter only if the team actively ran TensorZero's evaluations or fine-tuning workflows.
| Criterion | What to check | Why it matters after TensorZero |
|---|---|---|
| Performance | Overhead per request at sustained load | TensorZero users chose it partly for Rust-level latency |
| Provider coverage | Number of providers, OpenAI-compatible API | Migration is simplest when the same clients keep working |
| Routing and failover | Weighted routing, fallbacks, rule-based routing | Replaces TensorZero's routing, retries, and fallbacks |
| Governance | Virtual keys, budgets, rate limits, RBAC, SSO | Fills the gap TensorZero left around team-level control |
| MCP support | Tool discovery, filtering, per-user auth | Agents now call tools as often as they call models |
| Deployment | Self-hosted, Kubernetes, in-VPC, air-gapped | Regulated teams cannot route traffic through a third party |
| Maintenance | Active releases and a support path | The core reason to migrate in the first place |
The LLM Gateway Buyer's Guide expands each criterion into a scoring checklist. Teams that relied heavily on TensorZero's fine-tuning should plan for a separate optimization tool regardless of which gateway they pick, because none of the gateways below fine-tune models.
TensorZero Alternatives Compared at a Glance
The five TensorZero alternatives below are all LLM gateways, but they differ in deployment model and governance depth. Bifrost and LiteLLM are self-hosted open-source gateways, Kong AI Gateway and Agent Router extend existing API and service proxies, and Cloudflare AI Gateway is a managed service.
| Gateway | Deployment | Governance | MCP | Experimentation | Best fit |
|---|---|---|---|---|---|
| Bifrost | Self-hosted (Docker, NPX, Helm); Enterprise in-VPC, on-prem, air-gapped | Virtual keys, hierarchical budgets, rate limits, RBAC, OIDC and SCIM | MCP client and server, per-key tool filtering, Code Mode | Weighted routing, CEL routing rules, fallbacks | Enterprise platform teams |
| LiteLLM | Self-hosted proxy; hosted and enterprise tiers | Budgets per virtual key and user, rate limits | MCP gateway with key, team, and org permissions | Traffic mirroring for A/B tests | Python-centric teams |
| Kong AI Gateway | Konnect control plane with a local or self-managed data plane | Consumer groups, token budgets, rate limiting | LLM, MCP, and A2A traffic | Not published | Teams already on Kong |
| Agent Router (formerly Envoy AI Gateway) | Kubernetes-native on Envoy; local CLI; hosted options | Token limits per team, app, or model; upstream auth | Identity-filtered MCP tool catalog | Not published | Envoy and Kubernetes shops |
| Cloudflare AI Gateway | Managed service | Rate limiting, analytics, logging | Not published | Retries and model fallback | Teams on Cloudflare |
For a wider field, the comparison of top open-source LLM gateways covers additional self-hosted options.
1. Bifrost: Open Source AI Gateway for Governance and MCP
Bifrost is an open source AI gateway that unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, with governance, MCP, and enterprise deployment in the same binary. It adds 11 microseconds of overhead per request at 5,000 RPS with a 100% success rate in sustained benchmarks.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
Routing and reliability
Bifrost replaces TensorZero's routing, retries, and fallbacks with three layers. Automatic fallbacks switch to backup providers or models when the primary returns errors. Governance routing on a virtual key applies weighted load balancing across providers, and routing rules evaluate CEL expressions against headers, parameters, and capacity metrics before provider selection. Semantic caching supports exact-match and embedding-based replay, including streaming responses.
Governance and access control
Virtual keys are the primary governance entity in Bifrost. Each key carries allowed providers and models, budgets, and rate limits, and keys roll up into a customer, team, and virtual key hierarchy with independent budgets and limits at each level. Budgets can reset daily, weekly, monthly, quarterly, or yearly, and can align to calendar boundaries in UTC.
Bifrost Enterprise adds role-based access control, OIDC login and inbound SCIM 2.0 through user provisioning, and access profiles that issue governed virtual keys per user automatically. The governance resource page maps these controls to common policy requirements.

Figure 2: Access, spend, and routing policy run inside the request path, so a rejected or rerouted call never depends on application code.
MCP gateway
Bifrost acts as both an MCP client and an MCP server. It connects to external tool servers over STDIO, HTTP, or SSE and exposes those tools to clients such as Claude Desktop and Cursor through Bifrost as an MCP gateway. Tool calls are not auto-executed by default.
MCP tool filtering sets a strict allow-list per virtual key, and MCP authentication supports headers, OAuth 2.0, and per-user credentials. Code Mode lets the model write Python to orchestrate tools, cutting input tokens by up to 92.8% in large MCP deployments. The MCP gateway resource page covers the full design.
Enterprise deployment
Open-source Bifrost runs from NPX, Docker, or the official Helm chart, and a single instance handles roughly 3,000 to 5,000 RPS. Bifrost Enterprise adds clustering with gossip-based state sync and zero-downtime rolling updates, in-VPC deployments on AWS, GCP, Azure, Cloudflare, and Vercel, and on-prem or air-gapped installs. Guardrails, audit logs of administrative activity, and log exports to S3 or GCS round out compliance.
The Bifrost Enterprise page lists what changes from the open-source build.
What Bifrost does not replace: Bifrost does not fine-tune models or run automated prompt optimization. Teams that depended on TensorZero's fine-tuning loop should keep that workflow in a dedicated tool and point it at Bifrost logs exported over OpenTelemetry.
2. LiteLLM: Self-Hosted Proxy for Python Teams
LiteLLM is an open-source AI gateway and proxy that calls 100+ LLMs through an OpenAI-format API and tracks spend with budgets per virtual key or user. It suits teams already working in the Python ecosystem who want a self-hosted proxy with broad provider coverage and a large community.
LiteLLM's live docs list these gateway capabilities:
- Budgets and rate limits per virtual key and user, with spend tracking
- Guardrails, caching, and secret manager integrations as configurable proxy features
- An MCP gateway with a fixed endpoint for MCP tools, permission management by key, team, and organization, and Streamable HTTP, SSE, and stdio transports
- A/B testing through traffic mirroring, which sends production traffic to a silent secondary model for comparison
Best for: Python-centric engineering teams that want a self-hosted, open-source proxy with wide provider coverage and an established community.
Teams weighing this option against Bifrost can compare the two in the Bifrost LiteLLM alternative overview and the roundup of LiteLLM alternatives for 2026.
3. Kong AI Gateway: AI Traffic on an Existing API Platform
Kong AI Gateway is the AI layer of the Kong API platform, configured through the Konnect control plane and positioned as a single control point for LLM, MCP, and agent-to-agent (A2A) traffic. It suits organizations that already run Kong for API management and want AI traffic governed by the same team and tooling.
Kong's documentation describes these AI capabilities:
- Routing and load balancing across providers including OpenAI, Anthropic, Azure AI, Amazon Bedrock, and Gemini
- AI Consumer Groups that scope model access and token budgets by team or department
- AI Semantic Cache and AI Prompt Compressor to reduce token spend on repeated or long prompts
- Metering and billing for teams that resell AI access in tiers
The quickstart creates an AI Gateway control plane in Konnect and deploys a local data plane with Docker, with licensing handled through Konnect.
Best for: Platform teams standardized on Kong that want AI routing and budgets inside their existing API gateway estate.
Teams evaluating Kong against purpose-built gateways can review this list of Kong AI Gateway alternatives.
4. Agent Router (Formerly Envoy AI Gateway): Kubernetes-Native Routing
Agent Router is the new name for Envoy AI Gateway, now an Agentic AI Foundation project built on Envoy Proxy. It turns provider, credential, tool, and limit configuration into Envoy configuration, which suits Kubernetes teams that already operate Envoy for service traffic.
According to its project site, Agent Router provides:
- One OpenAI-compatible API routed to Anthropic, Bedrock, Vertex, Azure, or self-hosted vLLM, with 17 providers supported by default
- Traffic management including provider fallback, model name virtualization, and token limits per team, app, or model
- Upstream authentication so provider credentials stay with the router rather than in application workloads
- An MCP tool catalog aggregated from many MCP servers and filtered by the calling identity
- Inference-aware routing for self-managed models through Kubernetes InferencePool support
Observability follows the OpenTelemetry GenAI semantic conventions. A local CLI runs the router on a laptop, and the same configuration ships to a dedicated gateway or a Kubernetes cluster.
Best for: Kubernetes and Envoy shops that want AI routing expressed as Kubernetes resources and are comfortable operating Envoy.
More options in this category appear in the list of Envoy AI Gateway alternatives for LLM routing.
5. Cloudflare AI Gateway: Managed Gateway on Cloudflare
Cloudflare AI Gateway is a managed service that adds analytics, logging, caching, rate limiting, request retries, and model fallback in front of AI providers. It is available on all Cloudflare plans and needs only an endpoint change, which suits teams that want visibility without operating gateway infrastructure.
Cloudflare's documentation lists these features:
- Analytics for requests, tokens, and cost
- Logging of requests and errors
- Caching served from Cloudflare instead of the model provider
- Rate limiting, request retry, and model fallback for resilience
- Provider support including Workers AI, Anthropic, Google Gemini, OpenAI, and Replicate
Self-hosting, in-VPC deployment, and MCP governance are not published for Cloudflare AI Gateway, so the service fits teams whose data policies allow AI traffic to pass through a third-party network.
Best for: Teams already on Cloudflare that want managed observability and caching for LLM calls with minimal setup.
Teams that need self-hosted control instead can compare Cloudflare AI Gateway alternatives for enterprises.
Migrating from TensorZero to Bifrost
Migrating from TensorZero to Bifrost is mostly a client-side configuration change, because both gateways accept OpenAI-compatible requests. Teams update the client's base_url and model names, recreate routing and fallback settings as virtual keys and routing rules, and move rate limits into Bifrost budgets.
TensorZero clients call http://localhost:3000/openai/v1 with model names such as tensorzero::model_name::anthropic::claude-sonnet-4-6. With Bifrost as a drop-in replacement, the same OpenAI SDK points at the Bifrost endpoint and uses provider/model names, with a Bifrost virtual key as the API key:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8080/openai",
api_key="sk-bf-your-virtual-key",
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-4-6",
messages=[{"role": "user", "content": "Summarize this incident report."}],
)

Figure 3: Both gateways accept OpenAI-compatible requests, so the core switch is a base URL and model-name change rather than an application rewrite.
| TensorZero concept | Bifrost equivalent |
|---|---|
| Gateway API keys (Postgres-backed) | Virtual keys sent in the Authorization or x-bf-vk header |
| Custom rate limits by tag | Budgets and rate limits per customer, team, virtual key, and provider config |
| Routing, retries, fallbacks | Weighted governance routing, CEL routing rules, automatic fallbacks |
| Inference caching | Semantic caching with direct and similarity modes |
| Postgres or ClickHouse observability | Built-in observability on SQLite or Postgres, plus OpenTelemetry and Prometheus |
| Helm chart | Official Bifrost Helm chart and Terraform modules |
| Fine-tuning, evaluations, Autopilot | Not part of Bifrost; keep in a dedicated tool |
The enterprise deployment resource covers sizing and rollout for teams moving production traffic. Run both gateways in parallel during cutover, shift traffic by service, and retire TensorZero once error rates and spend match.
How to Choose the Best LLM Gateway for Your Team
The best LLM gateway to replace TensorZero depends on governance needs and deployment constraints more than on feature counts. Teams that need enterprise governance and MCP control should start with Bifrost; teams with existing platform commitments to Kong, Envoy, or Cloudflare can extend those instead.

Figure 4: Governance and deployment requirements narrow the field faster than feature lists do.
Three questions settle most decisions:
- Does the team need per-team budgets, SSO, RBAC, and MCP tool governance? Bifrost covers all four in one gateway, from the open-source build through Enterprise.
- Must traffic stay inside the company network? Self-hosted options are Bifrost, LiteLLM, Kong's locally deployed data plane, and Agent Router; Cloudflare AI Gateway is managed.
- Is there an existing proxy platform to extend? Kong and Envoy users can add AI routing to tooling they already operate, at the cost of depending on that platform's AI roadmap.
For a broader shortlist, the definitive guide to top LLM gateways ranks production options across managed and self-hosted deployments. Teams consolidating agent tool access should also read how an MCP gateway centralizes tool access.
Frequently Asked Questions
What is TensorZero?
TensorZero is an open-source LLMOps stack that combines a Rust LLM gateway with observability, evaluations, optimization, and experimentation. It stores inferences and feedback in Postgres or ClickHouse and uses that data for fine-tuning, prompt optimization, and adaptive A/B tests. The TensorZero homepage now states that the project remains available on GitHub but is no longer maintained, which is why many teams are evaluating replacement LLM gateways.
Is TensorZero still maintained?
No. As of October 2026, the TensorZero homepage states that TensorZero remains available on GitHub but is no longer maintained, and its docs URL redirects to the repository. Existing deployments keep running, but they will not receive provider updates, security fixes, or support. Teams running the TensorZero gateway in production should plan a migration to an actively maintained gateway such as Bifrost.
What is an LLM gateway?
An LLM gateway is a service that sits between applications and model providers and exposes one API for routing, failover, authentication, cost control, and logging. It removes provider-specific code from applications and centralizes policy. A thin proxy only forwards and translates requests; a governed gateway also enforces who can call which model, at what cost, and with which tools.
What is the best TensorZero alternative for enterprises?
Bifrost is the best TensorZero alternative for enterprises because it combines low gateway overhead (11 microseconds per request at 5,000 RPS) with virtual keys, hierarchical budgets, RBAC, SSO, MCP tool governance, and in-VPC or air-gapped deployment. Teams standardized on Kong or Envoy may extend those platforms instead, and Cloudflare users can adopt its managed gateway for lighter needs.
Can Bifrost replace TensorZero's A/B testing and fine-tuning?
Bifrost replaces TensorZero's gateway functions, including routing, fallbacks, caching, rate limits, and observability, and supports weighted traffic splits across providers and models. Bifrost does not fine-tune models or run automated prompt optimization. Teams that relied on those workflows should keep them in a dedicated tool and feed it with Bifrost request logs exported through OpenTelemetry trace export or log exports.
Is there an open source AI gateway that supports MCP?
Yes. Bifrost is an open source AI gateway that acts as both an MCP client and server, filters tools per virtual key, supports OAuth and per-user credentials, and offers Code Mode for large tool catalogs. LiteLLM and Agent Router also publish MCP gateway features. The Model Context Protocol defines how these tools are discovered, and Bifrost MCP gateway docs cover setup.
Try Bifrost Today
Bifrost is the most direct TensorZero replacement for teams that used TensorZero as their LLM gateway and now need a maintained, governed, enterprise-ready one. It keeps the OpenAI-compatible interface, adds virtual keys, budgets, MCP governance, and private deployment, and adds 11 microseconds of overhead per request at 5,000 RPS. To plan a TensorZero migration with the team, book a Bifrost demo or explore the Bifrost resources hub.