Try Bifrost Enterprise free for 14 days. Request access

Top 5 AI Gateways for Routing Between vLLM, Ollama, and Cloud LLMs

An AI gateway gives self-hosted inference servers and cloud LLM providers one API, one routing layer, and one place to enforce policy. This guide compares Bifrost, LiteLLM, Kong AI Gateway, Envoy AI Gateway, and Cloudflare AI Gateway for hybrid vLLM, Ollama, and cloud deployments.

Top 5 AI Gateways for Routing Between vLLM, Ollama, and Cloud LLMs

TL;DR

  • An AI gateway for hybrid deployments puts self-hosted servers (vLLM, Ollama, SGLang) and cloud providers behind one OpenAI-compatible endpoint, so moving a model between local GPUs and a cloud API is a configuration change.
  • Bifrost ships native vllm/, ollama/, and sgl/ providers alongside 25+ providers and 10,000+ models, with request-level fallbacks from self-hosted models to cloud models.
  • LiteLLM and Kong AI Gateway also name vLLM and Ollama as providers; Envoy AI Gateway and Cloudflare AI Gateway reach self-hosted servers through generic OpenAI-schema or custom-provider configuration.
  • Data residency in a hybrid setup is enforced by restricting which providers a caller may reach, not by trusting application code to pick the right model.
  • Bifrost adds 11 microseconds of overhead per request at 5,000 RPS in sustained benchmarks, which matters when the local inference path is supposed to be the fast one.

Teams running self-hosted inference servers next to OpenAI, Anthropic, or Bedrock often end up with two client stacks and no shared failover path, which is the problem an AI gateway solves. Bifrost, the open-source AI gateway written in Go and built by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability across self-hosted and cloud models. This guide compares five AI gateways on routing between vLLM, Ollama, SGLang, and cloud LLMs: native provider support, local-to-cloud fallback, cost shifting to owned GPUs, and data residency.

Why Hybrid LLM Deployments Need an AI Gateway

An AI gateway is a routing and policy layer that exposes one API to applications and forwards each request to the right model backend. In a hybrid deployment, those backends include self-hosted servers such as vLLM and Ollama alongside cloud providers, so the gateway becomes the single place where backend choice, failover, budgets, and logging are decided.

Without a gateway, each application hardcodes its own base URLs, so when a vLLM cluster is drained for an upgrade, nothing fails over. For background on the category, see this explainer on AI gateway architecture and why it matters.

Both vLLM and Ollama already speak the OpenAI wire format: vLLM ships an OpenAI-compatible server, and Ollama documents OpenAI API compatibility for a subset of endpoints. That compatibility is what makes a single gateway practical.

The Bifrost AI gateway translates one request format into each backend and lists every backend, local or cloud, in one supported providers matrix.

Three applications send requests to the Bifrost AI gateway, which routes each call to self-hosted vLLM, SGLang, or Ollama servers or to cloud LLM providers

Figure 1: Applications integrate once; moving a model between local GPUs and a cloud provider becomes a gateway configuration change.

The practical reasons teams put a gateway in front of a self-hosted LLM fleet:

  • Cost shifting: route high-volume, low-complexity traffic to owned GPUs and keep frontier cloud models for the requests that need them.
  • Resilience: fall back to a cloud model when a local server returns 5xx errors or is down for maintenance.
  • Data residency: pin regulated traffic to private LLM servers with no path to a public API.
  • One client contract: applications keep using the OpenAI SDK and change only the base URL with a drop-in replacement.

How to Evaluate an AI Gateway for vLLM, Ollama, and Cloud Routing

A hybrid AI gateway should be judged on four things: whether self-hosted servers are first-class providers, whether a request can fall back from a local model to a cloud model, whether traffic can be restricted to on-premise LLM servers, and where the gateway runs relative to your GPUs.

Criterion What to check Why it matters for hybrid routing
Native self-hosted providers Named vLLM, Ollama, and SGLang providers, not only a generic "OpenAI-compatible" slot Native providers handle per-server quirks such as base URLs, optional auth, and unsupported parameters
Local-to-cloud fallback Fallback chain that crosses providers and model names A local Llama model and a cloud GPT model never share a model ID
Load balancing across servers Weighted distribution across several inference servers of the same type One vLLM server per GPU node is common
Residency controls Ability to block cloud providers for specific callers or request classes Keeps restricted data on infrastructure you control
Deployment location Self-hosted binary or container, in-VPC, or vendor-hosted only A hosted-only gateway needs a public HTTPS path back to your GPUs
Overhead Per-request latency added by the gateway Local inference is often chosen for latency; the gateway should not erase that gain

A hosted-only gateway forces local servers onto the public internet, and one without cross-provider fallback cannot express "try vLLM, then OpenAI." The self-hosted side alone is covered in this roundup of LLM gateways for vLLM, SGLang, and Ollama, and the fallback mechanics are covered in the retries and fallbacks reference.

Self-Hosted LLM Support Compared at a Glance

The five gateways differ most on whether vLLM, Ollama, and SGLang are named providers and on where the gateway runs. Bifrost, LiteLLM, and Kong name vLLM and Ollama directly; Envoy AI Gateway uses OpenAI-schema backends; Cloudflare AI Gateway uses HTTPS custom providers.

Gateway vLLM Ollama SGLang Local-to-cloud fallback Where the gateway runs
Bifrost gateway Native vllm/ provider Native ollama/ provider Native sgl/ provider Request-level fallbacks and routing-rule fallbacks across providers Self-hosted (Docker, Kubernetes), in-VPC, on-prem
LiteLLM hosted_vllm/ provider ollama/ and ollama_chat/ Not published Fallbacks between model groups Self-hosted Python proxy
Kong AI Gateway vLLM provider (text generation) Ollama provider (chat, embeddings) Not published Priority balancer groups and failover targets Konnect control plane; quickstart runs a local Docker data plane
Envoy AI Gateway Self-hosted models via OpenAI schema Not published Not published Prioritized backendRefs with retry policy Kubernetes on Envoy Gateway; local aigw run
Cloudflare AI Gateway Custom provider (HTTPS base URL) Custom provider (HTTPS base URL) Custom provider (HTTPS base URL) Universal endpoint fallbacks, dynamic routes Cloudflare's network

"Not published" means the vendor documentation reviewed did not name that server; a generic OpenAI-compatible connection may still work, with per-server behavior left to you. Teams that need the gateway inside their own network should also read this guide to the best self-hosted AI gateway options.

1. Bifrost

Bifrost is an open-source AI gateway that treats vLLM, Ollama, and SGLang as native providers next to 25+ cloud and model providers, behind one OpenAI-compatible API. Routing, fallbacks, budgets, and logging apply to local and cloud backends alike, and Bifrost adds 11 microseconds of overhead per request at 5,000 RPS in published benchmarks.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

Native vLLM, Ollama, and SGLang providers

Each self-hosted server type has its own provider, addressed by a prefix such as vllm/meta-llama/Llama-3.1-8B-Instruct or ollama/llama3.2. The vLLM provider covers chat completions, the native Responses API, text completions, embeddings, rerank, transcription, and token counting, and it normalizes vLLM error payloads returned with HTTP 200 into standard errors.

The Ollama provider supports chat, Responses, text completions, embeddings, and model listing, with optional authentication. The SGLang provider covers chat, Responses, text completions, embeddings, and token counting. Any other OpenAI-compatible server can be added as a custom provider with its own base URL and an allowlist of request types.

Load balancing across inference servers

Every vLLM key carries its own server URL in vllm_key_config.url, so several keys with weights spread traffic across several vLLM servers. Bifrost selects keys by weighted random load balancing, and per-key model allowlists keep a request away from a server that has not loaded the requested model.

{
  "providers": {
    "vllm": {
      "keys": [
        {
          "name": "vllm-gpu-a",
          "value": "",
          "models": ["meta-llama/Llama-3.1-8B-Instruct"],
          "weight": 0.5,
          "vllm_key_config": {
            "url": "<http://vllm-a.internal:8000>",
            "model_name": "meta-llama/Llama-3.1-8B-Instruct"
          }
        },
        {
          "name": "vllm-gpu-b",
          "value": "",
          "models": ["meta-llama/Llama-3.1-8B-Instruct"],
          "weight": 0.5,
          "vllm_key_config": {
            "url": "<http://vllm-b.internal:8000>",
            "model_name": "meta-llama/Llama-3.1-8B-Instruct"
          }
        }
      ],
      "network_config": { "max_retries": 2 }
    },
    "ollama": {
      "keys": [
        {
          "name": "ollama-edge",
          "value": "",
          "models": ["*"],
          "weight": 1.0,
          "ollama_key_config": { "url": "<http://ollama.internal:11434>" }
        }
      ]
    }
  }
}

For multi-turn agents, session affinity keeps a conversation on the provider and key that served it before; because each vLLM key maps to one server, the session stays on one inference server. Bifrost Enterprise adds adaptive load balancing, which adjusts key weights from live error rates and latency and uses a circuit breaker to remove poorly performing keys from rotation.

Fallback from self-hosted models to cloud models

A request names its primary model and an ordered fallbacks list. If the vLLM provider still fails after its retry budget, Bifrost moves to the next entry, and the response's extra_fields.provider shows which backend answered. Cross-provider fallback is explicit, because a local model ID and a cloud model ID never match:

curl -X POST <http://localhost:8080/v1/chat/completions> \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vllm/meta-llama/Llama-3.1-8B-Instruct",
    "messages": [{"role": "user", "content": "Summarize this support ticket."}],
    "fallbacks": ["openai/gpt-4o-mini"]
  }'
An application request passes through Bifrost virtual keys, routing rules, and weighted key selection, which spreads traffic across two vLLM servers, an Ollama host, and a cloud provider

Figure 2: Virtual keys limit which backends a caller may reach, routing rules pick the target, and weighted keys spread load across individual inference servers.

Routing rules, virtual keys, and residency controls

As Figure 2 shows, three control points decide where a request lands. Virtual keys are deny-by-default: a key whose provider configs list only vllm and ollama cannot reach a cloud API, whatever model name the caller sends. Routing rules evaluate CEL expressions on headers, team, budget usage, and rate-limit usage, then pick weighted targets and an ordered fallback chain.

The complexity router classifies each request as SIMPLE, MEDIUM, or COMPLEX by embedding similarity, so a rule such as complexity_tier == "SIMPLE" can send short prompts to a local model while complex reasoning goes to a frontier model. Budgets sit on virtual keys, teams, and customers, and rate limits on virtual keys and provider configs, which makes the governance controls the same for local and cloud traffic.

Deployment, observability, and enterprise options

Bifrost runs as one container beside your inference servers, with built-in request logging, Prometheus metrics, and OpenTelemetry export.

For regulated environments, Bifrost supports in-VPC deployments, on-prem installs, and clustering for high availability. Teams evaluating those options can start from the Bifrost Enterprise page.

2. LiteLLM Proxy

LiteLLM is an open-source Python SDK and proxy server that routes to a long list of providers, including vLLM through the hosted_vllm/ prefix and Ollama through ollama/ and ollama_chat/. It is a common first choice for teams already working in Python who want a configurable proxy in front of mixed local and cloud models.

LiteLLM's vLLM documentation lists chat completions, embeddings, completions, rerank, and audio transcriptions, configured with an api_base per model. Its router groups deployments under a shared model_name, with weighted shuffle as the default strategy and latency-based, least-busy, usage-based, and cost-based routing as options.

Hybrid routing fit:

  • Fallbacks: defined from one model group to another, for example from a local Llama group to a cloud GPT group, with separate fallback types for context-window and content-policy errors.
  • SGLang: not named in the provider documentation reviewed; it would be configured as a generic OpenAI-compatible endpoint.

Best for: Python-centric teams that want a self-hosted proxy with broad provider coverage and are comfortable operating a Python service in the request path. Teams comparing performance and governance depth can review this LiteLLM migration resource.

3. Kong AI Gateway

Kong AI Gateway extends Kong's API gateway with AI providers, load balancing, and AI-specific policies, and its provider list includes both Ollama and vLLM. It fits organizations that already run Kong for API traffic and want hybrid LLM routing managed from the same Konnect control plane.

Kong's Ollama provider documents text generation and embeddings, and its vLLM provider documents text generation. Both are AI Model Provider entities with an upstream_url pointing at a self-hosted endpoint. The AI Gateway 2.0 provider pages for Ollama and vLLM are tagged "Incompatible with on-prem", so teams planning a fully on-prem Kong deployment should confirm support for their version.

Hybrid routing fit:

  • Load balancing algorithms: weighted round-robin, consistent hashing, least-connections, lowest-usage, lowest-latency, semantic, and priority.
  • Local-first with cloud fallback: the priority algorithm sends all traffic to the highest-priority group and falls back to lower groups only when every target in it is unavailable.
  • SGLang: not listed as a provider.

Best for: platform teams standardized on Kong who want AI routing inside existing API governance. A comparison of Kong alternatives for self-hosted AI gateways covers the trade-offs in more depth.

4. Envoy AI Gateway

Envoy AI Gateway, now renamed Agent Router under the Agentic AI Foundation, is an open-source, Kubernetes-native gateway built on Envoy Proxy and Envoy Gateway. Its stated goals include resilient connectivity across LLM providers and self-hosted models, and its docs describe using hosted and self-managed models within a single platform boundary.

Self-hosted models are configured with the OpenAI API schema, with support depending on the schema the server speaks; the docs give vLLM as an example. Ollama and SGLang are not named. For in-cluster inference, Envoy AI Gateway supports Gateway API InferencePool resources and custom endpoint picker providers for metric-aware endpoint selection.

Hybrid routing fit:

  • Provider fallback: a route lists prioritized backendRefs; the first is primary, and a retry policy in BackendTrafficPolicy triggers failover to the next healthy backend, including across cloud and on-premise providers.
  • Usage-based rate limiting: token-aware limits are configured through the gateway's policy APIs.

Best for: Kubernetes teams already running Envoy Gateway. This roundup of Envoy AI Gateway alternatives for LLM routing compares it with gateways that do not require Kubernetes.

5. Cloudflare AI Gateway

Cloudflare AI Gateway is a hosted gateway on Cloudflare's network that adds caching, rate limiting, observability, fallbacks, and dynamic routing in front of supported providers. Self-hosted models are reached through custom providers, which Cloudflare lists for connecting internal, self-hosted AI models.

A custom provider requires an HTTPS base URL that "must start with https://". In a hybrid setup, each vLLM or Ollama server therefore needs an HTTPS endpoint reachable from Cloudflare rather than a private address on an internal network.

Hybrid routing fit:

  • Fallbacks: the Universal endpoint accepts an ordered array of provider requests and returns the first success, with a response header that identifies which step answered.
  • Dynamic routing: visual or JSON flows with conditional, percentage, rate-limit, and budget-limit nodes, each able to switch to a fallback model.
  • Residency trade-off: because traffic transits a hosted service, keeping restricted prompts entirely on your own network is not possible with this design.

Best for: teams already on Cloudflare whose self-hosted models are exposed over HTTPS and who prefer a managed service to operating a gateway. This list of Cloudflare AI Gateway alternatives covers self-hosted options for teams that need the gateway inside their network.

Hybrid Model Routing Patterns for Self-Hosted and Cloud LLMs

Most hybrid deployments combine four model routing patterns: local-first with cloud fallback, cost shifting by request complexity, burst overflow to the cloud when local capacity is exhausted, and residency pinning that never leaves your network. Each pattern maps to a specific gateway mechanism, and mixing them on one gateway is the reason to centralize routing.

Pattern Goal How it is expressed in Bifrost
Local-first, cloud fallback Keep serving when a vLLM node fails or is drained Primary vllm/ model with openai/ or anthropic/ entries in fallbacks
Cost shifting by complexity Send simple prompts to owned GPUs Routing rule on complexity_tier == "SIMPLE" targeting an ollama/ or vllm/ model
Burst overflow Spill to cloud when local quotas are near their limit Routing rule on request > 90 or tokens_used > 90 with a cloud target
Residency pinning Restricted data never reaches a public API Virtual key allowing only self-hosted providers, plus a routing rule with an empty fallback list
A request passes routing rules to vLLM replicas, falls back to a cloud provider on 5xx errors, while restricted requests stay on self-hosted models

Figure 3: The fallback chain is a per-route decision: general traffic can spill to the cloud, while routes for restricted data are configured with an empty fallback list.

Residency pinning needs the most care, because one misconfigured fallback sends regulated text to a cloud API. In the Bifrost gateway, the reliable control is virtual key governance: issue restricted workloads a key whose provider configs include only self-hosted providers, so a cloud model is rejected even if a request asks for one. A routing rule can then direct classified traffic explicitly:

{
  "name": "restricted-stays-local",
  "enabled": true,
  "cel_expression": "headers[\"x-data-class\"] == \"restricted\"",
  "targets": [
    { "provider": "vllm", "model": "meta-llama/Llama-3.1-8B-Instruct", "weight": 1 }
  ],
  "fallbacks": [],
  "scope": "global",
  "priority": 0
}

Pair that with budgets and rate limits per virtual key to cap overflow cloud spend per team. The same pattern extends to fully disconnected sites, covered in this guide to air-gapped and on-prem AI gateways, and the broader concepts are in the AI gateway architecture explainer.

Frequently Asked Questions

What is an AI gateway?

An AI gateway is a routing and policy layer between applications and model backends that exposes one API and forwards each request to the right provider. For hybrid deployments, it unifies self-hosted servers such as vLLM and Ollama with cloud providers, and centralizes fallback, load balancing, budgets, and logging. The open-source Bifrost gateway provides this through an OpenAI-compatible API across 25+ providers and 10,000+ models.

Which LLM gateway is the best for routing between self-hosted and cloud models?

Bifrost is the strongest fit when self-hosted servers must be first-class providers: it has native vLLM, Ollama, and SGLang providers, cross-provider fallbacks, deny-by-default virtual keys for residency, and 11 microseconds of overhead at 5,000 RPS. LiteLLM suits Python-first teams, Kong suits existing Kong users, Envoy AI Gateway suits Kubernetes platform teams, and Cloudflare suits teams that accept a hosted gateway.

Is vLLM better than Ollama?

vLLM and Ollama serve different jobs. vLLM targets high-throughput GPU serving with features such as continuous batching, so it usually carries production traffic. Ollama prioritizes simple local setup and is common on developer machines and edge hosts. Many teams run both behind one gateway, using vLLM for production routes and Ollama for development or low-volume workloads.

What is the difference between SGLang and vLLM?

SGLang and vLLM are both open-source, high-throughput inference engines that expose OpenAI-compatible APIs, and both are used for GPU serving in production. They differ in scheduling and caching internals, and benchmark results vary by model and workload. Behind an AI gateway the difference is mostly operational: Bifrost addresses them as sgl/ and vllm/ providers with the same routing and fallback controls.

What is a vLLM router?

A vLLM router distributes requests across many vLLM workers in one deployment, with load balancing methods such as cache-aware and consistent hashing, and with support for prefill/decode disaggregation. It operates inside a single vLLM fleet. An AI gateway sits one layer above and routes across vLLM, Ollama, SGLang, and cloud providers, so the two are often deployed together.

Can Ollama connect to the internet?

Ollama runs models locally and binds to 127.0.0.1:11434 by default, so it only accepts connections from the same machine until OLLAMA_HOST is changed. Pulling models requires network access, and Ollama also offers cloud models through its own API. When Ollama sits behind a gateway, only the gateway needs network access to the Ollama host, and gateway access controls decide who can reach it.

Try Bifrost for Hybrid Self-Hosted and Cloud Routing

An AI gateway is what lets vLLM, Ollama, SGLang, and cloud LLMs operate as one fleet rather than separate integrations. Bifrost gives self-hosted servers native providers, routes between local and cloud models with explicit fallbacks, and enforces residency through virtual keys, all at 11 microseconds of overhead per request. To plan a hybrid deployment for your inference fleet, book a demo with the Bifrost team.