Contact Us

Self-Hosted LLM Gateways Compared: LiteLLM and Its Alternatives (2026)

Sep 22, 202612 min read
Origins AI banner: Self-Hosted LLM Gateways Compared: LiteLLM and Its Alternatives (2026)
openrouter alternatives litellm alternatives llm proxy llm gateway open source portkey alternatives

TL;DR

  • Check whether SSO, audit logs and RBAC sit in a paid enterprise tier before you promise them to security.
  • Issue gateway keys per team or tool and set budgets before the first developer connects, so spend is attributable.
  • A self-hosted gateway keeps prompts inside your network only while every upstream model is local.

Quick Answer: The self-hosted LiteLLM alternatives are Agent Router (formerly Envoy AI Gateway), Portkey's open-source gateway and Kong AI Gateway. Pick LiteLLM for built-in per-key budgets, Agent Router for Kubernetes-native routing under Apache 2.0, and Kong if you already run Kong. Cloudflare AI Gateway and OpenRouter are hosted services, not self-hosted.

The real decision isn't which project has the longest feature list. It's where spend controls, admin login and audit logs live, and whether those sit in the free core or a paid tier. LiteLLM documents per-key budgets in its open-source proxy but lists audit logs and fuller RBAC as enterprise features. Kong keeps token rate limiting and multi-model load balancing in its enterprise offering.

Also note that two well-known names aren't self-hosted: Cloudflare AI Gateway runs on Cloudflare's network, and OpenRouter is a hosted API.

Which self-hosted LLM gateways can an enterprise deploy in 2026?

Four projects come up most often when a team needs its LLM gateway open-source and self-hosted: LiteLLM, Agent Router (the project formerly called Envoy AI Gateway), Portkey's open-source gateway and Kong AI Gateway. All four run on your own infrastructure and accept OpenAI-style requests, but they come from different traditions:

How do LiteLLM, Envoy AI Gateway and other open-source gateways compare?

They differ most on where spend control lives. LiteLLM documents budgets per virtual key in its open-source proxy. Agent Router documents token quota policies and provider fallback. Kong places AI rate limiting and load balancing in its enterprise tier.

Capabilities as documented by each vendor on 21 September 2026; links in the text.

Gateway Runs inside your network Open-source core OpenAI-compatible endpoint Per-key budgets or token quotas Fallback across models or providers Admin SSO
LiteLLM Yes Yes (MIT outside enterprise/) Yes Yes Yes Yes (up to five users); Enterprise tier beyond
Agent Router (formerly Envoy AI Gateway) Yes (Kubernetes) Yes (Apache 2.0) Yes Yes (token quota policy) Yes Not documented
Portkey open-source gateway Yes Yes (MIT) Yes Not documented Yes Not documented
Kong AI Gateway Yes Yes (Kong Gateway, Apache 2.0) Yes Enterprise tier Enterprise tier Enterprise tier
Cloudflare AI Gateway No (Cloudflare's network) Not documented Yes Yes (spend limits, beta) Yes Not documented
OpenRouter No (hosted API) Not documented Yes Yes (per-key credit limits) Yes Enterprise tier

"Not documented" means we couldn't find the capability on the vendor's own pages on that date, not that it's missing. On Kong, the AI Rate Limiting Advanced plugin is available only in the AI Gateway Enterprise offering, as are AI Proxy Advanced and AI Semantic Cache.

Two practical notes. Palo Alto Networks acquired Portkey and now sells its commercial platform as Prisma AIRS AI Gateway (generally available July 2026). The open-source gateway is still published as Portkey-AI/gateway under the MIT license. On the commercial hybrid deployment, the control plane is hosted by the vendor while LLM traffic stays in your VPC, which some security reviews accept and some don't.

When each is the better fit. Choose LiteLLM when you want built-in per-key budgets and a large provider catalogue without a license. Choose Agent Router when your platform team already runs Envoy or Gateway API on Kubernetes. Choose Kong when Kong already fronts your other APIs. Choose Cloudflare AI Gateway or OpenRouter when a hosted endpoint is acceptable and you'd rather not run anything.

How do you deploy an internally hosted gateway for an engineering team?

Deploy it as an internal service behind your SSO, with one OpenAI-compatible endpoint, per-team keys and budgets, and logs shipped to your existing stack. The LLM proxy is small; the policy around it is the work.

A typical rollout for an engineering team runs in this order:

  1. Pick the deployment target. A Kubernetes namespace in your own cloud account or data center, with no public ingress.
  2. Put provider keys in a secret manager. Developers never see the OpenAI or Anthropic key; they get a gateway key.
  3. Issue keys per team or per tool. One key for the IDE plugin, one for CI, one per internal app, so spend is attributable.
  4. Set budgets and rate limits before the first developer connects, not after the first surprise invoice.
  5. Point the tools at the gateway through a custom OpenAI base URL.
  6. Wire logs and metrics into what the security and platform teams already watch.

The deployment checklist below is what a security reviewer will ask about.

Area What to decide A sensible default
Auth Who can call the gateway and who can administer it Gateway keys for callers; SSO for the admin UI (a paid tier on some projects)
Keys Where provider credentials live Secret manager; rotate on a schedule; never in client configs
Budgets Per team, per key or per model Per-team budget plus a per-key ceiling for CI jobs
Logging Metadata only, or prompts and responses too Metadata everywhere; full bodies only where policy allows, with retention limits
High availability Replicas, shared state, failover At least two gateway replicas behind an internal load balancer, shared Redis, managed Postgres
Upgrade path How versions and schema changes roll out Pin image versions; upgrade in staging first; check database migrations before production

How does a gateway sit in front of vLLM or other local model servers?

The gateway treats a local model server as one more OpenAI-compatible upstream. vLLM runs as a server that implements the OpenAI API protocol, started with vllm serve and optionally protected by --api-key, so the gateway routes to it exactly as it would to a hosted provider.

That design gives you three useful patterns:

This is also why teams looking for OpenRouter alternatives often end up self-hosting. They like one endpoint for many models; they don't want a hosted intermediary in the path.

Build it yourself or have it deployed for you: what changes?

Building it yourself costs engineering time rather than license fees; having it deployed moves the integration and on-call work to a partner. The gateway binary is the easy part. The work is SSO, key issuance, budgets, log pipelines, upgrades and the on-call rotation.

Build it yourself when:

Have it deployed for you when:

Most teams weighing Portkey alternatives or LiteLLM alternatives settle this ownership question before they pick a project. If you're building, DevOps and platform engineering support matters more than the choice between two MIT-licensed projects.

What does it take to run LiteLLM alternatives in production?

It takes horizontal replicas, a shared cache, a database for keys and spend, alerting, and a written upgrade process. LiteLLM's own guidance is concrete: give each worker pod 1 vCPU and 4Gi of memory and run Redis 7.0 or newer once you run more than one instance.

The same page recommends one worker per pod, a dedicated pod for scheduled jobs such as budget resets, a master key and quiet logs. It's a useful checklist whichever project you pick.

What you monitor matters as much as what you deploy:

A self-hosted LLM proxy is a stateful, security-relevant service. Run it with the same discipline as your identity provider. Teams without platform engineers on staff often bring one in, and top DevOps consulting companies in the US covers that market.

What mistakes should you avoid when self-hosting a model gateway?

The costly mistakes are about scope and ownership, not about which project you picked.

How Origins AI deploys and runs a self-hosted gateway for engineering teams

Origins AI (originshq.com) deploys a self-hosted gateway as part of the Origins AI Coding Tool. According to its product page, the gateway exposes an OpenAI-compatible REST API and proxies every LLM request from IDE plugins, CI pipelines and internal tools, with routing, rate limits, cost tracking and access control.

The deployment modes on the Coding Tool product page are on-premise, private cloud in your own AWS, Azure or GCP account, air-gapped with local models such as Llama, Mistral and CodeLlama, and hybrid. In on-premise and air-gapped modes, no source code is sent to an external service. In hybrid mode, the code context submitted to a hosted model leaves your network.

The page also lists secrets and PII filtering before content reaches the model, per-team RBAC and quotas, and local logging of every request. The same platform adds codebase intelligence and an AI code audit server for CI/CD, so the gateway isn't a standalone proxy. The company reports the gateway live in 1 week and the full platform in 30 days.

The same gateway layer sits under Origins AI Chat AI, which routes between OpenAI, Anthropic, Meta Llama, Mistral or your own fine-tuned model. The full catalogue is on the products overview. It's deployed by an implementation team inside your environment, not sold as a hosted subscription.

Talk to an engineer

If you're weighing a self-hosted gateway for your engineering team and want to walk through deployment modes, identity and audit with an engineer, book a technical call.

Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.

Frequently Asked Questions

What is the best alternative to LiteLLM?
It depends on your platform. Agent Router suits teams that already run Envoy or Kubernetes Gateway API, and it keeps the Apache 2.0 license. Kong AI Gateway suits teams that already run Kong, though several AI plugins need its enterprise offering. Portkey's open-source gateway starts with a single npx or Docker command.
Is LiteLLM production ready for an enterprise?
Its docs publish explicit production guidance: horizontal pods, Redis, Postgres, a master key and a separate job worker. The enterprise questions are about features. SSO is free for up to five users. Basic roles (proxy admin, internal user) are in the open-source proxy; organizations and fuller RBAC need Enterprise.
Is there anything better than OpenRouter for a company?
For a company that needs prompts to stay in its own network, a self-hosted gateway fits better, because OpenRouter is a hosted API. OpenRouter is still a good choice for prototypes and for teams comfortable with a hosted intermediary. It offers per-key credit limits and automatic fallbacks, which covers basic spend control for small teams.
What has to be monitored on a self-hosted gateway?
Watch four signals: latency the gateway adds on top of the model, error rates per upstream provider, spend per key against budget, and rejected requests from rate limits or guardrails. Add alerts on replica health, on Redis and on the database that stores keys and spend, and review guardrail blocks weekly for false positives. Track upgrade drift too: an unpatched gateway sits in front of every model. A silent gateway failure looks like every AI tool breaking at once.
Do developers have to change their editors or SDKs?
Usually not. Many IDE plugins, command-line tools and SDKs accept a custom OpenAI-compatible base URL and key, so developers swap the provider key for a gateway key. Tools that hard-code one vendor endpoint are the exception, so test each one first.
Can a gateway restrict which models each team may use?
Yes, on most projects. Gateways typically attach allowed models, budgets and rate limits to a key or a team, so a contractor key can be limited to one local model while a platform team gets hosted models too. Check whether per-team model access sits in the open-source core or an enterprise tier on the project you choose.
Book a call

About the Author

Apoorva Kumar is Co-Founder and CEO of Origins AI (originshq.com), an AI engineering partner for product teams building AI workflows, AI agents and LLM integrations. A CSE graduate of IIT Kharagpur, Apoorva previously built and scaled technology at Sony, NuCash, YesMadam and FrontPage.