Quick Answer: The self-hosted LiteLLM alternatives are Agent Router (formerly Envoy AI Gateway), Portkey's open-source gateway and Kong AI Gateway. Pick LiteLLM for built-in per-key budgets, Agent Router for Kubernetes-native routing under Apache 2.0, and Kong if you already run Kong. Cloudflare AI Gateway and OpenRouter are hosted services, not self-hosted.
The real decision isn't which project has the longest feature list. It's where spend controls, admin login and audit logs live, and whether those sit in the free core or a paid tier. LiteLLM documents per-key budgets in its open-source proxy but lists audit logs and fuller RBAC as enterprise features. Kong keeps token rate limiting and multi-model load balancing in its enterprise offering.
Also note that two well-known names aren't self-hosted: Cloudflare AI Gateway runs on Cloudflare's network, and OpenRouter is a hosted API.
Which self-hosted LLM gateways can an enterprise deploy in 2026?
Four projects come up most often when a team needs its LLM gateway open-source and self-hosted: LiteLLM, Agent Router (the project formerly called Envoy AI Gateway), Portkey's open-source gateway and Kong AI Gateway. All four run on your own infrastructure and accept OpenAI-style requests, but they come from different traditions:
- LiteLLM is a Python proxy built around virtual keys, budgets and a large provider catalogue. It ships as a Docker image with Kubernetes and Helm options.
- Agent Router is built on Envoy and runs on Kubernetes. The maintainers renamed it from Envoy AI Gateway and moved it to the Agentic AI Foundation, keeping the Apache 2.0 license and the existing CRD names, API group and CLI.
- Portkey's gateway is a lightweight MIT-licensed router you can start with
npxor Docker. The commercial edition, now Prisma AIRS AI Gateway, is a hybrid: the data plane runs in your VPC while the vendor hosts the control plane. - Kong AI Gateway adds AI plugins to Kong Gateway, whose core is Apache 2.0. Data planes run self-hosted or connect to Kong's Konnect control plane.
How do LiteLLM, Envoy AI Gateway and other open-source gateways compare?
They differ most on where spend control lives. LiteLLM documents budgets per virtual key in its open-source proxy. Agent Router documents token quota policies and provider fallback. Kong places AI rate limiting and load balancing in its enterprise tier.
Capabilities as documented by each vendor on 21 September 2026; links in the text.
| Gateway | Runs inside your network | Open-source core | OpenAI-compatible endpoint | Per-key budgets or token quotas | Fallback across models or providers | Admin SSO |
|---|---|---|---|---|---|---|
| LiteLLM | Yes | Yes (MIT outside enterprise/) |
Yes | Yes | Yes | Yes (up to five users); Enterprise tier beyond |
| Agent Router (formerly Envoy AI Gateway) | Yes (Kubernetes) | Yes (Apache 2.0) | Yes | Yes (token quota policy) | Yes | Not documented |
| Portkey open-source gateway | Yes | Yes (MIT) | Yes | Not documented | Yes | Not documented |
| Kong AI Gateway | Yes | Yes (Kong Gateway, Apache 2.0) | Yes | Enterprise tier | Enterprise tier | Enterprise tier |
| Cloudflare AI Gateway | No (Cloudflare's network) | Not documented | Yes | Yes (spend limits, beta) | Yes | Not documented |
| OpenRouter | No (hosted API) | Not documented | Yes | Yes (per-key credit limits) | Yes | Enterprise tier |
"Not documented" means we couldn't find the capability on the vendor's own pages on that date, not that it's missing. On Kong, the AI Rate Limiting Advanced plugin is available only in the AI Gateway Enterprise offering, as are AI Proxy Advanced and AI Semantic Cache.
Two practical notes. Palo Alto Networks acquired Portkey and now sells its commercial platform as Prisma AIRS AI Gateway (generally available July 2026). The open-source gateway is still published as Portkey-AI/gateway under the MIT license. On the commercial hybrid deployment, the control plane is hosted by the vendor while LLM traffic stays in your VPC, which some security reviews accept and some don't.
When each is the better fit. Choose LiteLLM when you want built-in per-key budgets and a large provider catalogue without a license. Choose Agent Router when your platform team already runs Envoy or Gateway API on Kubernetes. Choose Kong when Kong already fronts your other APIs. Choose Cloudflare AI Gateway or OpenRouter when a hosted endpoint is acceptable and you'd rather not run anything.
How do you deploy an internally hosted gateway for an engineering team?
Deploy it as an internal service behind your SSO, with one OpenAI-compatible endpoint, per-team keys and budgets, and logs shipped to your existing stack. The LLM proxy is small; the policy around it is the work.
A typical rollout for an engineering team runs in this order:
- Pick the deployment target. A Kubernetes namespace in your own cloud account or data center, with no public ingress.
- Put provider keys in a secret manager. Developers never see the OpenAI or Anthropic key; they get a gateway key.
- Issue keys per team or per tool. One key for the IDE plugin, one for CI, one per internal app, so spend is attributable.
- Set budgets and rate limits before the first developer connects, not after the first surprise invoice.
- Point the tools at the gateway through a custom OpenAI base URL.
- Wire logs and metrics into what the security and platform teams already watch.
The deployment checklist below is what a security reviewer will ask about.
| Area | What to decide | A sensible default |
|---|---|---|
| Auth | Who can call the gateway and who can administer it | Gateway keys for callers; SSO for the admin UI (a paid tier on some projects) |
| Keys | Where provider credentials live | Secret manager; rotate on a schedule; never in client configs |
| Budgets | Per team, per key or per model | Per-team budget plus a per-key ceiling for CI jobs |
| Logging | Metadata only, or prompts and responses too | Metadata everywhere; full bodies only where policy allows, with retention limits |
| High availability | Replicas, shared state, failover | At least two gateway replicas behind an internal load balancer, shared Redis, managed Postgres |
| Upgrade path | How versions and schema changes roll out | Pin image versions; upgrade in staging first; check database migrations before production |
How does a gateway sit in front of vLLM or other local model servers?
The gateway treats a local model server as one more OpenAI-compatible upstream. vLLM runs as a server that implements the OpenAI API protocol, started with vllm serve and optionally protected by --api-key, so the gateway routes to it exactly as it would to a hosted provider.
That design gives you three useful patterns:
- Mixed routing. Send code completion to a local model on vLLM and harder reasoning tasks to a hosted model, under one endpoint and one set of keys.
- Model-aware load balancing on Kubernetes. The Kubernetes Gateway API Inference Extension is an official Kubernetes project for routing to self-hosted models by model name, including LoRA variants. Agent Router supports its InferencePool resource, and the project names Envoy Gateway, kgateway and GKE Gateway as implementations.
- Fully offline operation. If every upstream is local, no prompt has to leave your network. That only holds while every route stays local.
This is also why teams looking for OpenRouter alternatives often end up self-hosting. They like one endpoint for many models; they don't want a hosted intermediary in the path.
Build it yourself or have it deployed for you: what changes?
Building it yourself costs engineering time rather than license fees; having it deployed moves the integration and on-call work to a partner. The gateway binary is the easy part. The work is SSO, key issuance, budgets, log pipelines, upgrades and the on-call rotation.
Build it yourself when:
- You have a platform team that already runs Kubernetes, Redis and Postgres in production.
- One open-source project covers your controls without a paid tier.
- You're comfortable owning upgrades and security patches.
Have it deployed for you when:
- The security review needs air-gapped operation, secrets and PII filtering, or audit trails beyond the free tier.
- You need the gateway tied into CI/CD, code review or internal apps, not just a proxy.
- Nobody on the team wants the pager for another stateful service.
Most teams weighing Portkey alternatives or LiteLLM alternatives settle this ownership question before they pick a project. If you're building, DevOps and platform engineering support matters more than the choice between two MIT-licensed projects.
What does it take to run LiteLLM alternatives in production?
It takes horizontal replicas, a shared cache, a database for keys and spend, alerting, and a written upgrade process. LiteLLM's own guidance is concrete: give each worker pod 1 vCPU and 4Gi of memory and run Redis 7.0 or newer once you run more than one instance.
The same page recommends one worker per pod, a dedicated pod for scheduled jobs such as budget resets, a master key and quiet logs. It's a useful checklist whichever project you pick.
What you monitor matters as much as what you deploy:
- Latency added by the gateway, separate from model latency, so you can tell which one is slow.
- Error rates per upstream, since a provider outage should trigger fallback, not a flood of failed IDE requests.
- Spend per key and per team, against budget, with alerts before the limit.
- Rejected requests, from rate limits, guardrails or expired keys, which usually signal a misconfigured tool.
A self-hosted LLM proxy is a stateful, security-relevant service. Run it with the same discipline as your identity provider. Teams without platform engineers on staff often bring one in, and top DevOps consulting companies in the US covers that market.
What mistakes should you avoid when self-hosting a model gateway?
The costly mistakes are about scope and ownership, not about which project you picked.
- Assuming a self-hosted gateway keeps prompts private. Requests routed to a hosted model still go out of your network. Only local upstreams keep prompts inside.
- Logging full prompts by default. Prompts contain source code, customer data and secrets. Decide retention and redaction first.
- Skipping per-key budgets. One runaway CI loop or agent can consume a month's allowance overnight.
- Running a single replica. When the gateway is down, every AI tool your engineers use is down.
- Missing the paid-tier boundary. Check whether SSO, audit logs or RBAC sit in an enterprise tier before you promise them to security.
- Letting provider keys leak into client configs. The point of the gateway is that developers never hold provider credentials.
How Origins AI deploys and runs a self-hosted gateway for engineering teams
Origins AI (originshq.com) deploys a self-hosted gateway as part of the Origins AI Coding Tool. According to its product page, the gateway exposes an OpenAI-compatible REST API and proxies every LLM request from IDE plugins, CI pipelines and internal tools, with routing, rate limits, cost tracking and access control.
The deployment modes on the Coding Tool product page are on-premise, private cloud in your own AWS, Azure or GCP account, air-gapped with local models such as Llama, Mistral and CodeLlama, and hybrid. In on-premise and air-gapped modes, no source code is sent to an external service. In hybrid mode, the code context submitted to a hosted model leaves your network.
The page also lists secrets and PII filtering before content reaches the model, per-team RBAC and quotas, and local logging of every request. The same platform adds codebase intelligence and an AI code audit server for CI/CD, so the gateway isn't a standalone proxy. The company reports the gateway live in 1 week and the full platform in 30 days.
The same gateway layer sits under Origins AI Chat AI, which routes between OpenAI, Anthropic, Meta Llama, Mistral or your own fine-tuned model. The full catalogue is on the products overview. It's deployed by an implementation team inside your environment, not sold as a hosted subscription.
Talk to an engineer
If you're weighing a self-hosted gateway for your engineering team and want to walk through deployment modes, identity and audit with an engineer, book a technical call.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


