Quick Answer: An LLM gateway is one endpoint between your apps and every model provider that routes requests, issues keys with budgets and applies guardrails. Tools call the gateway instead of holding provider keys. Enterprise options include LiteLLM, Portkey, Kong AI Gateway and Azure API Management, which differ mostly on self-hosting and which controls need an enterprise license.
Most teams meet the problem before the term. Three coding assistants, two chat pilots and a support bot each hold their own provider key, and finance can't tell who spent what. An LLM gateway puts one governed endpoint in front of every model.
A security reviewer asks four things: where it runs, who holds provider keys, what is logged, and which tier holds the controls.
What is an LLM gateway and what does it do?
An LLM gateway is a reverse proxy built for model traffic. Your IDE plugins, CI jobs and internal apps send requests to one OpenAI-compatible or Anthropic-format endpoint, and the gateway decides which model serves each one. It does three jobs:
- Routing. It picks the provider and model for each request, balances load across keys and deployments, and falls back when one fails. LLM routing can follow cost, task type or team policy.
- Keys and budgets. Provider keys stay server-side. Each developer or team gets a scoped gateway credential with its own rate limits, token quotas or spend budget.
- Guardrails and audit. It can redact PII, block prompt injection and log every request and token count for review.
Microsoft and Kong call the same layer an AI gateway, and both extend it to MCP servers and agent APIs. A narrower "LLM router" only chooses the model; a gateway also governs who may call it.
What are the best enterprise LLM gateway solutions for managing internal AI coding tools?
The best choice depends on where your models already live and whether the gateway must run inside your network. Four options are worth comparing first:
- LiteLLM, a gateway you run yourself from its public code, with virtual keys, spend tracking and per-team budgets in its docs.
- Portkey, a gateway with a managed control plane. SSO, RBAC and budget limits sit in its Enterprise offering, which also adds a hybrid mode with the data plane in your cloud.
- Kong AI Gateway, AI policies on top of Kong's API gateway, with data plane nodes that run in your environment.
- Azure API Management, whose AI features also back the gateway inside Microsoft Foundry.
For coding tools, first check that the assistant accepts a custom base URL. Claude Code does, through ANTHROPIC_BASE_URL, and Anthropic's gateway guide for Claude Code lists credentials, usage attribution, budgets and audit logging as the reasons to use one.
How do LiteLLM, Portkey, Kong AI Gateway and Azure AI Foundry compare?
Microsoft has renamed Azure AI Foundry to Microsoft Foundry, and its gateway (in preview) runs on Azure API Management's AI gateway capabilities, so the Azure column describes that service.
| Capability | LiteLLM | Portkey | Kong AI Gateway | Azure API Management |
|---|---|---|---|---|
| Multi-model routing and fallbacks | Yes | Yes | Yes | Yes |
| Scoped credentials per team or user | Yes | Yes | Yes | Yes |
| Spend or token budgets per team | Yes | Enterprise tier | Enterprise tier | Yes |
| Guardrails on prompts | Yes | Yes | Yes | Yes |
| Self-hosted data plane | Yes | Yes | Yes | Yes |
| SSO and RBAC | Enterprise tier | Enterprise tier | Enterprise tier | Yes |
Capabilities as documented by each vendor on 21 September 2026; links in the text.
LiteLLM's SSO is free for up to five users and needs an enterprise license beyond that. Basic roles (proxy admin, internal user) are in the free proxy; organizations and fuller RBAC need Enterprise. Portkey has replaced virtual keys with a Model Catalog, its budget limits are on the Enterprise plan and for select Pro customers, its guardrails come in tiers by plan, and its gateway package can run locally, while the hybrid mode with a managed control plane is an Enterprise feature. Kong documents team token budgets and cost limits through AI Consumer Groups, but both rely on the AI Rate Limiting Advanced plugin, sold only with its AI Gateway Enterprise offering. Azure's budgets are token quotas rather than spend caps, its guardrails run through Azure AI Content Safety, and its self-hosted gateway runs only on the Developer and Premium tiers. Microsoft documents that it needs outbound 443 to Azure, so it isn't built for a fully air-gapped network.
In short, choose LiteLLM when your platform team wants to run and patch the gateway itself. Choose Portkey when you want a managed control plane and can buy the Enterprise tier for SSO and budgets. Choose Kong when Kong already fronts your APIs. Choose Azure when your models already run in Azure and Foundry.
LLM gateway vs API gateway vs MCP gateway: what is the difference?
The difference is the unit each one governs. An API gateway meters requests; an LLM gateway meters tokens and models; an MCP gateway governs the tools that agents call.
| Gateway type | Governs | Typical controls |
|---|---|---|
| API gateway | HTTP requests to your services | Auth, rate limits per request, routing by path |
| LLM gateway | Model calls | Token quotas, spend budgets, model fallback, prompt guardrails |
| MCP gateway | Tool calls over Model Context Protocol | Tool registry, OAuth to tools, per-agent access |
Azure says its gateway features extend its existing API gateway and now cover MCP servers and agent APIs. Kong can turn existing APIs into MCP tools that agents discover and call. Treat an MCP gateway as a feature to check for, not a separate purchase.
When does an engineering team need a gateway at all?
You need one once more than one team calls more than one model and someone has to answer for the bill.
The usual triggers: a second provider arrives and apps hard-code two SDKs, a security review asks where prompts are logged, a departed engineer's key is still in a CI secret, or finance asks which team drove last month's token spike. Code-level routing can handle the first trigger; the other three need a shared gateway with identity and audit.
How does an LLM gateway control AI spend across teams?
It gives every caller its own credential and counts tokens against that credential before and after each call. Three mechanisms do the work:
- Attribution. Every request carries a team, project or developer key, so usage rolls up without log mining.
- Limits. Token rate limits stop spikes; quotas or spend budgets cap a period. Azure's token limit policy, for example, returns a 429 on a rate breach and a 403 when a quota runs out.
- Cheaper routing. Simple tasks go to a smaller model, and semantic caching reuses answers to near-identical prompts.
What mistakes should you avoid when rolling out an LLM gateway?
- Leaving direct keys in place. If developers keep personal provider keys, your budgets and logs cover only part of the traffic.
- Treating it as set-and-forget. Anthropic warns that a gateway which doesn't forward new Claude Code capabilities breaks those features, so plan upgrades.
- Buying the wrong tier. SSO, budgets and self-hosting are often enterprise features.
- Logging prompts without a retention policy. Full prompt logs are useful for audit and risky for PII. Decide what's kept, where, and for how long.
- Skipping the hybrid question. A self-hosted gateway calling a hosted model still sends the prompt out. Only local models keep the content inside.
If the gateway is part of a wider build, scope it alongside the AI engineering services that will consume it.
How Origins AI deploys an LLM gateway inside the customer's network
Origins AI (originshq.com) ships the gateway as part of the Origins AI Coding Tool, a self-hosted AI coding assistant and LLM gateway that its implementation team deploys in your environment. According to its product page, the gateway exposes an OpenAI-compatible REST API, routes by task type, cost or team policy, logs every request and token count locally, and filters secrets and PII before content reaches a model.
It runs in four modes: on-premise, private cloud in your own AWS, Azure or GCP account, air-gapped with local models such as Llama and Mistral, and hybrid. In on-premise and air-gapped modes, no source code is sent to an external service. Hybrid mode keeps the gateway local but sends the submitted context to the hosted model. Origins AI reports the gateway can be live in one week. Air-gapped deployments need a model running locally, and the best Ollama alternatives for running LLMs compares those runtimes.
Per its product page, Origins AI Chat AI has its own gateway and delivery layer with multi-model routing across OpenAI, Anthropic, Meta Llama, Mistral or a fine-tuned model. The wider catalogue is on the products page.
Talk to an engineer
Weighing where a gateway should run and who will operate it? Book a call with an Origins AI engineer to walk through your deployment mode and model mix.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


