Try Bifrost Enterprise free for 14 days. Request access

Top 4 LLM Gateways for Self-Hosted Models: vLLM, SGLang and Ollama (2026)

Compare the top LLM gateway options for self-hosted vLLM, SGLang, and Ollama models on replica load balancing, cloud failover, hybrid routing, and governance.

Top 4 LLM Gateways for Self-Hosted Models: vLLM, SGLang and Ollama (2026)

TL;DR

  • An LLM gateway in front of self-hosted models gives vLLM, SGLang, and Ollama servers one OpenAI-compatible endpoint, plus load balancing, retries, failover, access control, and request logs.
  • Bifrost ships native vLLM, SGLang, and Ollama providers, weighted load balancing across inference replicas, and fallbacks from self-hosted models to cloud providers, with 11 microseconds of overhead per request at 5,000 RPS.
  • LiteLLM reaches vLLM and Ollama through dedicated provider routes and offers several routing strategies across deployments that share a model name.
  • Kong lists vLLM and Ollama as AI providers, but its multi-target load balancer (AI Proxy Advanced) is part of the AI Gateway Enterprise offering.
  • Apache APISIX reaches self-hosted servers through a generic OpenAI-compatible provider and balances them with its ai-proxy-multi plugin.

Teams that run open-weight models on their own GPUs usually end up with several inference servers: a vLLM pool for production traffic, an SGLang deployment for high-throughput jobs, and Ollama on developer machines, each with its own URL, model names, and failure modes. An LLM gateway puts one API, one set of credentials, and one routing policy in front of all of them. Bifrost, the open-source AI gateway built in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, including hybrid estates that mix self-hosted and cloud models. This guide compares four LLM gateways on how well they front vLLM, SGLang, and Ollama specifically.

Why Self-Hosted LLM Deployments Need an LLM Gateway

A self-hosted LLM deployment needs an LLM gateway because inference servers serve models but do not govern traffic. vLLM, SGLang, and Ollama each expose an OpenAI-compatible API, yet none of them decides which replica gets a request, retries a failed call, enforces per-team access, or fails over to a cloud model when a GPU node restarts.

Engines like vLLM focus on GPU efficiency (its PagedAttention paper describes the memory management behind its throughput), and SGLang focuses on efficient execution of structured language model programs. The gateway handles what appears once more than one team or server is involved. For a broader primer on that layer, see this complete guide to LLM gateways for enterprise AI.

Layered stack with applications on top, an LLM gateway in the middle, and vLLM, SGLang, and Ollama inference servers plus a hosted cloud API at the bottom

Figure 1: Applications call one OpenAI-compatible endpoint; the gateway decides which inference server or cloud model serves each request.

As Figure 1 shows, the gateway turns many servers into one service. It absorbs these problems:

  • Replica sprawl: a vLLM pool with four GPU nodes has four URLs, and clients should not hard-code any of them.
  • Restarts and OOM events: model reloads and out-of-memory crashes return 5xx errors or refused connections that every client would otherwise retry on its own.
  • Mixed estates: production often pairs a self-hosted LLM with a hosted model for overflow or for tasks the open-weight model handles poorly.
  • No built-in tenancy: inference servers rarely ship per-team keys, budgets, or audit-friendly request logs.

This post is about the gateway in front of self-hosted models. If the question is which gateways you can self-host themselves, the companion piece on open-source LLM gateways for self-hosted deployments covers that angle.

Key Criteria for Evaluating an LLM Gateway for vLLM, SGLang, and Ollama

The right gateway for self-hosted models supports each inference engine natively, balances load across replicas, retries transient failures, fails over to another provider, routes by policy, and runs inside your network. Generic OpenAI-compatible proxying is a starting point, not the full requirement.

Criterion What to check Why it matters for self-hosted models
Engine support Named providers for vLLM, SGLang, Ollama, or only a generic OpenAI-compatible route Native providers handle engine quirks such as filtered parameters and non-standard error payloads
Replica load balancing Weighted or performance-based selection across several server URLs GPU nodes differ in capacity; equal round-robin overloads smaller nodes
Retries and failover Backoff on 5xx and connection errors, then a fallback provider Model reloads and node restarts are routine on self-managed GPUs
Hybrid routing Rules that send some traffic local and some to cloud models Keeps routine traffic on owned GPUs and sends only hard requests to paid APIs
Governance Per-team keys, model allowlists, budgets, rate limits Self-hosted servers usually have no tenancy of their own
Deployment In-VPC, on-prem, and air-gapped support; internal TLS Self-hosting is often a data-residency decision, so the gateway must stay inside too

The LLM Gateway Buyer's Guide covers each criterion in more depth.

LLM Gateways for Self-Hosted Models Compared at a Glance

The four gateways differ most on whether vLLM, SGLang, and Ollama are first-class providers and on how replica load balancing is packaged. Bifrost names all three engines in its provider catalog; LiteLLM and Kong name two each; Apache APISIX treats all three as generic OpenAI-compatible endpoints.

Capability Bifrost LiteLLM Kong AI Gateway Apache APISIX
vLLM Native provider hosted_vllm/ provider Listed provider openai-compatible with endpoint override
SGLang Native provider Generic OpenAI-compatible route Not published openai-compatible with endpoint override
Ollama Native provider ollama/ and ollama_chat/ providers Listed provider openai-compatible with endpoint override
Replica load balancing Weighted per server key; adaptive scoring in Enterprise Routing strategies across deployments sharing a model name Several algorithms in AI Proxy Advanced (AI Gateway Enterprise) Weighted round-robin, consistent hashing, semantic
Failover to cloud Request-level and routing-rule fallback chains Fallbacks, context-window and content-policy fallbacks Retries and failover between targets fallback_strategy on 429, 5xx, or health
Hybrid routing rules CEL rules on headers, team, budget, complexity tier Model-group fallbacks and routing strategies Priority, semantic, lowest-latency balancing Priority and semantic balancing
Overhead 11 µs per request at 5,000 RPS Not published Not published Not published

1. Bifrost

The Bifrost AI gateway is a high-performance, open-source gateway that unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, and it treats self-hosted inference servers as first-class providers rather than generic endpoints. vLLM, SGLang, and Ollama each have a dedicated provider with engine-specific handling.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

Native vLLM, SGLang, and Ollama providers

Each engine has its own provider page and configuration block, and the API key can stay blank for local servers:

  • vLLM: the vLLM provider covers chat and text completions, embeddings, rerank, transcription, and streaming, and sends Responses API calls natively to /v1/responses. An optional use_anthropic_endpoints setting routes chat through vLLM's Anthropic-compatible /v1/messages endpoint per key or per model alias. Bifrost also normalizes vLLM error payloads that arrive with HTTP 200 into standard errors.
  • SGLang: the SGLang provider supports chat, text completions, embeddings, tool calling, and streaming, and strips fields SGLang does not accept, such as store and service_tier.
  • Ollama: the Ollama provider points at http://localhost:11434 or a remote Ollama host and supports chat, embeddings, and tool calling.

For both vLLM and SGLang, Bifrost drops Anthropic-hosted server tools such as web_search from requests, because a self-hosted server cannot run them. That keeps clients with built-in web search enabled from failing every request. Any other OpenAI-compatible server can be added as a custom provider with its own base URL, path overrides, and a trusted CA certificate for internal or air-gapped endpoints.

Load balancing across inference replicas

In Bifrost, each vLLM key carries its own server URL and a weight, so a replica pool is simply a list of keys. Weighted load balancing picks a key per request by weighted random selection, which lets a node with twice the GPU memory take twice the traffic. Keys can also carry model allowlists, so one key serves only the 70B model while another serves the embedding model. Bifrost Enterprise adds adaptive load balancing, which rescores every route every five seconds on error rate and latency and circuit-breaks routes that keep failing.

Hybrid routing with routing rules

Routing rules evaluate CEL expressions on headers, virtual key, team, customer, budget usage, and rate-limit usage, then send the request to weighted targets with their own fallback chain. The Complexity Router adds a complexity_tier variable (SIMPLE, MEDIUM, or COMPLEX), assigned by embedding each request and matching it against 150 default reference phrases.

A request with a virtual key enters Bifrost, where a routing rule reads the complexity tier and sends traffic to Ollama, a vLLM pool, or a cloud model

Figure 2: Routing rules keep routine traffic on self-hosted GPUs and send only the hardest requests to a paid frontier model.

In Figure 2, simple and medium requests stay on owned hardware and only complex ones reach a hosted model. Virtual keys wrap this in governance, with provider allowlists, budgets, and rate limits per team, so a development key can be restricted to Ollama while production keys reach the vLLM pool and the cloud fallback.

Deployment and observability

Bifrost starts with a single npx command or a container image and supports in-VPC deployments for teams whose reason for self-hosting is data control.

Built-in observability records inputs, outputs, tokens, cost, and latency for every request asynchronously, and the published benchmarks show 11 microseconds of overhead per request at 5,000 RPS. Bifrost Enterprise adds clustering for high availability.

2. LiteLLM

LiteLLM is a Python SDK and proxy server that exposes many model providers through an OpenAI-style interface, and its documentation includes dedicated pages for both vLLM and Ollama.

What the documentation shows for self-hosted engines:

  • vLLM: reached through the hosted_vllm/ prefix with an api_base pointing at the server; supported endpoints include /chat/completions, /embeddings, /completions, /rerank, and /audio/transcriptions.
  • Ollama: reached through ollama/ or the recommended ollama_chat/ prefix, with streaming, JSON mode, and tool calling examples.
  • SGLang: no dedicated provider page is published; SGLang's OpenAI-compatible server can be called through the generic OpenAI-compatible route with the openai/ prefix and a custom api_base.

For load balancing, LiteLLM groups deployments that share a model_name and selects among them with a configurable routing_strategy, including simple-shuffle, least-busy, usage-based, latency-based, and cost-based routing. Reliability features include retries, cooldowns, ordered fallbacks, and dedicated context-window and content-policy fallbacks, with Redis recommended for tracking cooldowns and usage across instances in production.

Best for: Python-centric teams that want a broad provider catalog and SDK-level access to vLLM and Ollama, and that can operate Redis alongside the proxy for multi-instance state. Teams comparing it against higher-throughput options can review these LiteLLM alternatives for 2026.

3. Kong AI Gateway

Kong AI Gateway extends Kong Gateway with AI-specific plugins, and its provider documentation lists both vLLM and Ollama among supported AI providers.

The AI Proxy plugin translates requests to the configured provider format and documents fulfillment of requests to self-hosted models. Multi-target balancing lives in AI Proxy Advanced, which supports round-robin, consistent hashing, least-connections, lowest-latency, lowest-usage, semantic, and priority algorithms, plus retries, failover between targets, health checks, and a circuit breaker. Three details matter for self-hosted estates:

  • Licensing: AI Proxy Advanced is documented as available only as part of the AI Gateway Enterprise offering.
  • Deployment model: AI Gateway is managed through Konnect with data planes in your environment, and an on-prem configuration is documented; the vLLM and Ollama provider entities in AI Gateway 2.0 are marked incompatible with on-prem.
  • SGLang: no SGLang provider is listed in the current documentation.

Best for: enterprises already standardized on Kong Gateway that want AI traffic managed as another set of plugins, and that are licensed for AI Gateway Enterprise. Teams weighing a move away can compare Kong AI Gateway alternatives.

4. Apache APISIX

Apache APISIX is an Apache Software Foundation API gateway whose ai-proxy and ai-proxy-multi plugins add LLM routing. It does not name vLLM, SGLang, or Ollama as providers; instead, its openai-compatible provider forwards requests to any custom endpoint set in override.endpoint, which covers all three engines through their OpenAI-compatible APIs.

The ai-proxy-multi plugin is where self-hosted pools are managed:

  • Balancing: weighted round-robin, consistent hashing on headers, cookies, or consumer, and a semantic algorithm that matches prompts to instance examples.
  • Priority: an instance priority setting that takes precedence over weight, for local-first, cloud-second ordering.
  • Fallback: fallback_strategy values for instance health and rate limiting, HTTP 429, and HTTP 5xx, bounded by max_retries and retry_on_failure_within_ms.
  • Telemetry: access-log fields for token usage, model, and time to first response, consumed by the logging plugins.

The documentation notes that the semantic algorithm does not participate in health checks or fallback retries, so failures on a semantically chosen instance return to the client.

Best for: teams already running APISIX for API traffic that are comfortable assembling LLM routing from generic plugins. For purpose-built options, see these open-source AI gateways for self-hosted LLM deployments.

vLLM vs SGLang vs Ollama: What Changes at the Gateway

vLLM, SGLang, and Ollama all speak the OpenAI API, but they differ in which endpoints they expose, which makes engine-aware providers useful at the gateway. vLLM serves the widest set of operations, SGLang targets high-throughput serving, and Ollama is local-first, so each needs slightly different handling.

The table shows how Bifrost handles each engine, a useful checklist when testing any gateway.

Behavior in Bifrost vLLM SGLang Ollama
Default local URL in docs http://localhost:8000 http://localhost:8000 http://localhost:11434
Responses API Native /v1/responses Converted to chat completions Converted to chat completions
Anthropic Messages mode Optional, per key or alias Optional, per key or alias Not listed
Embeddings Supported Supported Supported
Rerank Supported Not listed Not listed
Transcription Supported Not supported Not supported

This is why vLLM vs Ollama is rarely either-or: Ollama suits laptops and edge boxes, vLLM or SGLang carry production load, and a gateway lets one client reach both. Ollama's own OpenAI compatibility documentation lists the endpoints it exposes for that purpose. For a worked example of routing a coding tool through Ollama, see this guide to a self-hosted AI gateway for Cursor with Claude or Ollama.

Hybrid LLM Routing: Self-Hosted First, Cloud as Fallback

Hybrid routing sends traffic to self-hosted models by default and to a cloud provider only when local capacity fails or a request needs a larger model. In Bifrost, this combines weighted replica selection, per-provider retries with exponential backoff, and an ordered fallback chain that can end at a hosted model.

A request is assigned to a weighted vLLM replica, transient errors are retried with exponential backoff, and after retries are exhausted Bifrost falls back to a cloud provider

Figure 3: Each provider in the chain gets its own retry budget, so a GPU node restart costs a few retries instead of a failed request.

Figure 3 traces the failure path. Retries and fallbacks in Bifrost work as two nested layers:

  1. Retries: on 5xx responses and network errors such as a refused connection, Bifrost retries the same provider with exponential backoff and jitter, starting at 500 ms and capped at 5 seconds by default; max_retries is set per provider and defaults to 0.
  2. Fallbacks: once retries are exhausted, Bifrost moves to the next provider/model in the request's fallbacks array, and each fallback receives its own full retry budget.
  3. Attribution: the response's extra_fields.provider field records which provider served the request, so dashboards can show how often traffic left your GPUs.

Each fallback runs as a fresh request, so governance and logging plugins apply to overflow traffic too, and chains can be attached to a routing rule instead of each request. For more on designing these chains, see this guide to automatic failover and load balancing for LLM apps, and for weighting strategies across many replicas, the comprehensive guide to load balancing in an AI gateway.

How to Choose an LLM Gateway for Self-Hosted Models

Choose a gateway for self-hosted models by matching it to your engine mix, existing infrastructure, and governance needs. Native engine support and hybrid routing matter most for production GPU estates; reusing an existing API gateway matters most when AI traffic is a small addition to established API operations.

  • Mixed vLLM, SGLang, and Ollama estate with cloud fallback: Bifrost covers all three engines natively and adds routing rules, virtual keys, and in-VPC deployment. The walkthrough on multi-provider routing and custom providers in Bifrost shows the configuration side.
  • Python-first prototyping on vLLM or Ollama: LiteLLM's SDK and provider routes are familiar to Python teams.
  • Existing Kong deployment with an enterprise license: Kong AI Gateway keeps AI and API policies in one place.
  • Existing APISIX deployment: APISIX's generic OpenAI-compatible provider works if engine-specific handling is not required.

Regulated teams should weigh deployment isolation first; this roundup of air-gapped and on-prem AI gateways for regulated industries compares that dimension directly. For a wider view of the production LLM gateway market beyond self-hosted models, this production-ready comparison of LLM gateways covers hosted-provider scenarios as well.

Frequently Asked Questions

What is an LLM gateway?

An LLM gateway is a service that sits between applications and model providers, exposing one API while handling routing, load balancing, retries, failover, authentication, and logging. For self-hosted models, it gives vLLM, SGLang, and Ollama servers one OpenAI-compatible endpoint, so operators can change backends without changing application code.

Which AI gateway is the best for self-hosted models?

Bifrost is the strongest fit for self-hosted models when teams run more than one inference engine or need cloud fallback. It has native vLLM, SGLang, and Ollama providers, weighted load balancing across replicas, CEL routing rules for hybrid traffic, and virtual keys for governance, with 11 microseconds of overhead per request at 5,000 RPS. Teams already committed to Kong or APISIX may prefer to extend those gateways instead.

How do I load balance multiple vLLM instances?

Put a gateway in front of the instances and register each vLLM server as a weighted backend. In Bifrost, each vLLM key holds its own server URL and weight, and weighted key selection distributes requests. Set max_retries on the provider so transient 5xx errors are retried, and add a fallback provider so requests still complete if the whole pool is unavailable.

Is vLLM better than Ollama?

vLLM and Ollama serve different jobs rather than one being better. vLLM is built for GPU-efficient, high-concurrency serving in production, while Ollama is a local-first engine for running models on personal computers or single servers. Many teams use Ollama for development and vLLM for production behind one gateway endpoint.

What is the difference between SGLang and vLLM?

SGLang and vLLM are both high-throughput open-source inference servers with OpenAI-compatible APIs. vLLM is known for PagedAttention memory management and exposes a broad endpoint set, including rerank, transcription, and a native Responses API. SGLang was designed for efficient execution of structured language model programs. Behind a gateway, both can be benchmarked side by side on real traffic.

Can a gateway fail over from a self-hosted LLM to a cloud provider?

Yes. Bifrost retries the self-hosted provider with exponential backoff on 5xx and connection errors, then moves to the next entry in the fallback chain, which can be a hosted model such as one served through AWS Bedrock or Azure OpenAI. Each fallback gets its own retry budget, and the response records which provider served it. This walkthrough of routing, fallback, and governance in Bifrost shows a full chain.

Try Bifrost Today

Self-hosted models give teams control over cost and data, and an LLM gateway is what makes a fleet of vLLM, SGLang, and Ollama servers behave like one reliable service. Bifrost adds native engine support, replica load balancing, hybrid routing rules, and cloud fallback in an open-source gateway that runs inside your own network, with governance controls for every team that uses it. To see how Bifrost can front your self-hosted LLM infrastructure, book a demo with the Bifrost team.