Top 4 LLM Gateways for Self-Hosted Models: vLLM, SGLang and Ollama (2026)
Compare the top LLM gateway options for self-hosted vLLM, SGLang, and Ollama models on replica load balancing, cloud failover, hybrid routing, and governance.
TL;DR
- An LLM gateway in front of self-hosted models gives vLLM, SGLang, and Ollama servers one OpenAI-compatible endpoint, plus load balancing, retries, failover, access control, and request logs.
- Bifrost ships native vLLM, SGLang, and Ollama providers, weighted load balancing across inference replicas, and fallbacks from self-hosted models to cloud providers, with 11 microseconds of overhead per request at 5,000 RPS.
- LiteLLM reaches vLLM and Ollama through dedicated provider routes and offers several routing strategies across deployments that share a model name.
- Kong lists vLLM and Ollama as AI providers, but its multi-target load balancer (AI Proxy Advanced) is part of the AI Gateway Enterprise offering.
- Apache APISIX reaches self-hosted servers through a generic OpenAI-compatible provider and balances them with its ai-proxy-multi plugin.
Teams that run open-weight models on their own GPUs usually end up with several inference servers: a vLLM pool for production traffic, an SGLang deployment for high-throughput jobs, and Ollama on developer machines, each with its own URL, model names, and failure modes. An LLM gateway puts one API, one set of credentials, and one routing policy in front of all of them. Bifrost, the open-source AI gateway built in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, including hybrid estates that mix self-hosted and cloud models. This guide compares four LLM gateways on how well they front vLLM, SGLang, and Ollama specifically.
Why Self-Hosted LLM Deployments Need an LLM Gateway
A self-hosted LLM deployment needs an LLM gateway because inference servers serve models but do not govern traffic. vLLM, SGLang, and Ollama each expose an OpenAI-compatible API, yet none of them decides which replica gets a request, retries a failed call, enforces per-team access, or fails over to a cloud model when a GPU node restarts.
Engines like vLLM focus on GPU efficiency (its PagedAttention paper describes the memory management behind its throughput), and SGLang focuses on efficient execution of structured language model programs. The gateway handles what appears once more than one team or server is involved. For a broader primer on that layer, see this complete guide to LLM gateways for enterprise AI.

Figure 1: Applications call one OpenAI-compatible endpoint; the gateway decides which inference server or cloud model serves each request.
As Figure 1 shows, the gateway turns many servers into one service. It absorbs these problems:
- Replica sprawl: a vLLM pool with four GPU nodes has four URLs, and clients should not hard-code any of them.
- Restarts and OOM events: model reloads and out-of-memory crashes return 5xx errors or refused connections that every client would otherwise retry on its own.
- Mixed estates: production often pairs a self-hosted LLM with a hosted model for overflow or for tasks the open-weight model handles poorly.
- No built-in tenancy: inference servers rarely ship per-team keys, budgets, or audit-friendly request logs.
This post is about the gateway in front of self-hosted models. If the question is which gateways you can self-host themselves, the companion piece on open-source LLM gateways for self-hosted deployments covers that angle.
Key Criteria for Evaluating an LLM Gateway for vLLM, SGLang, and Ollama
The right gateway for self-hosted models supports each inference engine natively, balances load across replicas, retries transient failures, fails over to another provider, routes by policy, and runs inside your network. Generic OpenAI-compatible proxying is a starting point, not the full requirement.
| Criterion | What to check | Why it matters for self-hosted models |
|---|---|---|
| Engine support | Named providers for vLLM, SGLang, Ollama, or only a generic OpenAI-compatible route | Native providers handle engine quirks such as filtered parameters and non-standard error payloads |
| Replica load balancing | Weighted or performance-based selection across several server URLs | GPU nodes differ in capacity; equal round-robin overloads smaller nodes |
| Retries and failover | Backoff on 5xx and connection errors, then a fallback provider | Model reloads and node restarts are routine on self-managed GPUs |
| Hybrid routing | Rules that send some traffic local and some to cloud models | Keeps routine traffic on owned GPUs and sends only hard requests to paid APIs |
| Governance | Per-team keys, model allowlists, budgets, rate limits | Self-hosted servers usually have no tenancy of their own |
| Deployment | In-VPC, on-prem, and air-gapped support; internal TLS | Self-hosting is often a data-residency decision, so the gateway must stay inside too |
The LLM Gateway Buyer's Guide covers each criterion in more depth.
LLM Gateways for Self-Hosted Models Compared at a Glance
The four gateways differ most on whether vLLM, SGLang, and Ollama are first-class providers and on how replica load balancing is packaged. Bifrost names all three engines in its provider catalog; LiteLLM and Kong name two each; Apache APISIX treats all three as generic OpenAI-compatible endpoints.
| Capability | Bifrost | LiteLLM | Kong AI Gateway | Apache APISIX |
|---|---|---|---|---|
| vLLM | Native provider | hosted_vllm/ provider |
Listed provider | openai-compatible with endpoint override |
| SGLang | Native provider | Generic OpenAI-compatible route | Not published | openai-compatible with endpoint override |
| Ollama | Native provider | ollama/ and ollama_chat/ providers |
Listed provider | openai-compatible with endpoint override |
| Replica load balancing | Weighted per server key; adaptive scoring in Enterprise | Routing strategies across deployments sharing a model name | Several algorithms in AI Proxy Advanced (AI Gateway Enterprise) | Weighted round-robin, consistent hashing, semantic |
| Failover to cloud | Request-level and routing-rule fallback chains | Fallbacks, context-window and content-policy fallbacks | Retries and failover between targets | fallback_strategy on 429, 5xx, or health |
| Hybrid routing rules | CEL rules on headers, team, budget, complexity tier | Model-group fallbacks and routing strategies | Priority, semantic, lowest-latency balancing | Priority and semantic balancing |
| Overhead | 11 µs per request at 5,000 RPS | Not published | Not published | Not published |
1. Bifrost
The Bifrost AI gateway is a high-performance, open-source gateway that unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, and it treats self-hosted inference servers as first-class providers rather than generic endpoints. vLLM, SGLang, and Ollama each have a dedicated provider with engine-specific handling.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
Native vLLM, SGLang, and Ollama providers
Each engine has its own provider page and configuration block, and the API key can stay blank for local servers:
- vLLM: the vLLM provider covers chat and text completions, embeddings, rerank, transcription, and streaming, and sends Responses API calls natively to
/v1/responses. An optionaluse_anthropic_endpointssetting routes chat through vLLM's Anthropic-compatible/v1/messagesendpoint per key or per model alias. Bifrost also normalizes vLLM error payloads that arrive with HTTP 200 into standard errors. - SGLang: the SGLang provider supports chat, text completions, embeddings, tool calling, and streaming, and strips fields SGLang does not accept, such as
storeandservice_tier. - Ollama: the Ollama provider points at
http://localhost:11434or a remote Ollama host and supports chat, embeddings, and tool calling.
For both vLLM and SGLang, Bifrost drops Anthropic-hosted server tools such as web_search from requests, because a self-hosted server cannot run them. That keeps clients with built-in web search enabled from failing every request. Any other OpenAI-compatible server can be added as a custom provider with its own base URL, path overrides, and a trusted CA certificate for internal or air-gapped endpoints.
Load balancing across inference replicas
In Bifrost, each vLLM key carries its own server URL and a weight, so a replica pool is simply a list of keys. Weighted load balancing picks a key per request by weighted random selection, which lets a node with twice the GPU memory take twice the traffic. Keys can also carry model allowlists, so one key serves only the 70B model while another serves the embedding model. Bifrost Enterprise adds adaptive load balancing, which rescores every route every five seconds on error rate and latency and circuit-breaks routes that keep failing.
Hybrid routing with routing rules
Routing rules evaluate CEL expressions on headers, virtual key, team, customer, budget usage, and rate-limit usage, then send the request to weighted targets with their own fallback chain. The Complexity Router adds a complexity_tier variable (SIMPLE, MEDIUM, or COMPLEX), assigned by embedding each request and matching it against 150 default reference phrases.

Figure 2: Routing rules keep routine traffic on self-hosted GPUs and send only the hardest requests to a paid frontier model.
In Figure 2, simple and medium requests stay on owned hardware and only complex ones reach a hosted model. Virtual keys wrap this in governance, with provider allowlists, budgets, and rate limits per team, so a development key can be restricted to Ollama while production keys reach the vLLM pool and the cloud fallback.
Deployment and observability
Bifrost starts with a single npx command or a container image and supports in-VPC deployments for teams whose reason for self-hosting is data control.
Built-in observability records inputs, outputs, tokens, cost, and latency for every request asynchronously, and the published benchmarks show 11 microseconds of overhead per request at 5,000 RPS. Bifrost Enterprise adds clustering for high availability.
2. LiteLLM
LiteLLM is a Python SDK and proxy server that exposes many model providers through an OpenAI-style interface, and its documentation includes dedicated pages for both vLLM and Ollama.
What the documentation shows for self-hosted engines:
- vLLM: reached through the
hosted_vllm/prefix with anapi_basepointing at the server; supported endpoints include/chat/completions,/embeddings,/completions,/rerank, and/audio/transcriptions. - Ollama: reached through
ollama/or the recommendedollama_chat/prefix, with streaming, JSON mode, and tool calling examples. - SGLang: no dedicated provider page is published; SGLang's OpenAI-compatible server can be called through the generic OpenAI-compatible route with the
openai/prefix and a customapi_base.
For load balancing, LiteLLM groups deployments that share a model_name and selects among them with a configurable routing_strategy, including simple-shuffle, least-busy, usage-based, latency-based, and cost-based routing. Reliability features include retries, cooldowns, ordered fallbacks, and dedicated context-window and content-policy fallbacks, with Redis recommended for tracking cooldowns and usage across instances in production.
Best for: Python-centric teams that want a broad provider catalog and SDK-level access to vLLM and Ollama, and that can operate Redis alongside the proxy for multi-instance state. Teams comparing it against higher-throughput options can review these LiteLLM alternatives for 2026.
3. Kong AI Gateway
Kong AI Gateway extends Kong Gateway with AI-specific plugins, and its provider documentation lists both vLLM and Ollama among supported AI providers.
The AI Proxy plugin translates requests to the configured provider format and documents fulfillment of requests to self-hosted models. Multi-target balancing lives in AI Proxy Advanced, which supports round-robin, consistent hashing, least-connections, lowest-latency, lowest-usage, semantic, and priority algorithms, plus retries, failover between targets, health checks, and a circuit breaker. Three details matter for self-hosted estates:
- Licensing: AI Proxy Advanced is documented as available only as part of the AI Gateway Enterprise offering.
- Deployment model: AI Gateway is managed through Konnect with data planes in your environment, and an on-prem configuration is documented; the vLLM and Ollama provider entities in AI Gateway 2.0 are marked incompatible with on-prem.
- SGLang: no SGLang provider is listed in the current documentation.
Best for: enterprises already standardized on Kong Gateway that want AI traffic managed as another set of plugins, and that are licensed for AI Gateway Enterprise. Teams weighing a move away can compare Kong AI Gateway alternatives.
4. Apache APISIX
Apache APISIX is an Apache Software Foundation API gateway whose ai-proxy and ai-proxy-multi plugins add LLM routing. It does not name vLLM, SGLang, or Ollama as providers; instead, its openai-compatible provider forwards requests to any custom endpoint set in override.endpoint, which covers all three engines through their OpenAI-compatible APIs.
The ai-proxy-multi plugin is where self-hosted pools are managed:
- Balancing: weighted round-robin, consistent hashing on headers, cookies, or consumer, and a semantic algorithm that matches prompts to instance examples.
- Priority: an instance priority setting that takes precedence over weight, for local-first, cloud-second ordering.
- Fallback:
fallback_strategyvalues for instance health and rate limiting, HTTP 429, and HTTP 5xx, bounded bymax_retriesandretry_on_failure_within_ms. - Telemetry: access-log fields for token usage, model, and time to first response, consumed by the logging plugins.
The documentation notes that the semantic algorithm does not participate in health checks or fallback retries, so failures on a semantically chosen instance return to the client.
Best for: teams already running APISIX for API traffic that are comfortable assembling LLM routing from generic plugins. For purpose-built options, see these open-source AI gateways for self-hosted LLM deployments.
vLLM vs SGLang vs Ollama: What Changes at the Gateway
vLLM, SGLang, and Ollama all speak the OpenAI API, but they differ in which endpoints they expose, which makes engine-aware providers useful at the gateway. vLLM serves the widest set of operations, SGLang targets high-throughput serving, and Ollama is local-first, so each needs slightly different handling.
The table shows how Bifrost handles each engine, a useful checklist when testing any gateway.
| Behavior in Bifrost | vLLM | SGLang | Ollama |
|---|---|---|---|
| Default local URL in docs | http://localhost:8000 |
http://localhost:8000 |
http://localhost:11434 |
| Responses API | Native /v1/responses |
Converted to chat completions | Converted to chat completions |
| Anthropic Messages mode | Optional, per key or alias | Optional, per key or alias | Not listed |
| Embeddings | Supported | Supported | Supported |
| Rerank | Supported | Not listed | Not listed |
| Transcription | Supported | Not supported | Not supported |
This is why vLLM vs Ollama is rarely either-or: Ollama suits laptops and edge boxes, vLLM or SGLang carry production load, and a gateway lets one client reach both. Ollama's own OpenAI compatibility documentation lists the endpoints it exposes for that purpose. For a worked example of routing a coding tool through Ollama, see this guide to a self-hosted AI gateway for Cursor with Claude or Ollama.
Hybrid LLM Routing: Self-Hosted First, Cloud as Fallback
Hybrid routing sends traffic to self-hosted models by default and to a cloud provider only when local capacity fails or a request needs a larger model. In Bifrost, this combines weighted replica selection, per-provider retries with exponential backoff, and an ordered fallback chain that can end at a hosted model.

Figure 3: Each provider in the chain gets its own retry budget, so a GPU node restart costs a few retries instead of a failed request.
Figure 3 traces the failure path. Retries and fallbacks in Bifrost work as two nested layers:
- Retries: on 5xx responses and network errors such as a refused connection, Bifrost retries the same provider with exponential backoff and jitter, starting at 500 ms and capped at 5 seconds by default;
max_retriesis set per provider and defaults to 0. - Fallbacks: once retries are exhausted, Bifrost moves to the next
provider/modelin the request'sfallbacksarray, and each fallback receives its own full retry budget. - Attribution: the response's
extra_fields.providerfield records which provider served the request, so dashboards can show how often traffic left your GPUs.
Each fallback runs as a fresh request, so governance and logging plugins apply to overflow traffic too, and chains can be attached to a routing rule instead of each request. For more on designing these chains, see this guide to automatic failover and load balancing for LLM apps, and for weighting strategies across many replicas, the comprehensive guide to load balancing in an AI gateway.
How to Choose an LLM Gateway for Self-Hosted Models
Choose a gateway for self-hosted models by matching it to your engine mix, existing infrastructure, and governance needs. Native engine support and hybrid routing matter most for production GPU estates; reusing an existing API gateway matters most when AI traffic is a small addition to established API operations.
- Mixed vLLM, SGLang, and Ollama estate with cloud fallback: Bifrost covers all three engines natively and adds routing rules, virtual keys, and in-VPC deployment. The walkthrough on multi-provider routing and custom providers in Bifrost shows the configuration side.
- Python-first prototyping on vLLM or Ollama: LiteLLM's SDK and provider routes are familiar to Python teams.
- Existing Kong deployment with an enterprise license: Kong AI Gateway keeps AI and API policies in one place.
- Existing APISIX deployment: APISIX's generic OpenAI-compatible provider works if engine-specific handling is not required.
Regulated teams should weigh deployment isolation first; this roundup of air-gapped and on-prem AI gateways for regulated industries compares that dimension directly. For a wider view of the production LLM gateway market beyond self-hosted models, this production-ready comparison of LLM gateways covers hosted-provider scenarios as well.
Frequently Asked Questions
What is an LLM gateway?
An LLM gateway is a service that sits between applications and model providers, exposing one API while handling routing, load balancing, retries, failover, authentication, and logging. For self-hosted models, it gives vLLM, SGLang, and Ollama servers one OpenAI-compatible endpoint, so operators can change backends without changing application code.
Which AI gateway is the best for self-hosted models?
Bifrost is the strongest fit for self-hosted models when teams run more than one inference engine or need cloud fallback. It has native vLLM, SGLang, and Ollama providers, weighted load balancing across replicas, CEL routing rules for hybrid traffic, and virtual keys for governance, with 11 microseconds of overhead per request at 5,000 RPS. Teams already committed to Kong or APISIX may prefer to extend those gateways instead.
How do I load balance multiple vLLM instances?
Put a gateway in front of the instances and register each vLLM server as a weighted backend. In Bifrost, each vLLM key holds its own server URL and weight, and weighted key selection distributes requests. Set max_retries on the provider so transient 5xx errors are retried, and add a fallback provider so requests still complete if the whole pool is unavailable.
Is vLLM better than Ollama?
vLLM and Ollama serve different jobs rather than one being better. vLLM is built for GPU-efficient, high-concurrency serving in production, while Ollama is a local-first engine for running models on personal computers or single servers. Many teams use Ollama for development and vLLM for production behind one gateway endpoint.
What is the difference between SGLang and vLLM?
SGLang and vLLM are both high-throughput open-source inference servers with OpenAI-compatible APIs. vLLM is known for PagedAttention memory management and exposes a broad endpoint set, including rerank, transcription, and a native Responses API. SGLang was designed for efficient execution of structured language model programs. Behind a gateway, both can be benchmarked side by side on real traffic.
Can a gateway fail over from a self-hosted LLM to a cloud provider?
Yes. Bifrost retries the self-hosted provider with exponential backoff on 5xx and connection errors, then moves to the next entry in the fallback chain, which can be a hosted model such as one served through AWS Bedrock or Azure OpenAI. Each fallback gets its own retry budget, and the response records which provider served it. This walkthrough of routing, fallback, and governance in Bifrost shows a full chain.
Try Bifrost Today
Self-hosted models give teams control over cost and data, and an LLM gateway is what makes a fleet of vLLM, SGLang, and Ollama servers behave like one reliable service. Bifrost adds native engine support, replica load balancing, hybrid routing rules, and cloud fallback in an open-source gateway that runs inside your own network, with governance controls for every team that uses it. To see how Bifrost can front your self-hosted LLM infrastructure, book a demo with the Bifrost team.