Last updated: 3 October 2026
Quick Answer: The best local LLM for coding is decided by VRAM: gpt-oss-20b at 16 GB, Devstral Small 2 24B at 24 to 32 GB. One 80 GB GPU opens gpt-oss-120b or a 4-bit Qwen3-Coder-Next, and DeepSeek-V4-Flash and GLM-5.3-Flash need a multi-GPU server. Take the largest model that fits, then test it on your own repository.
Most rankings of the best LLM for coding assume you can rent a GPU the size of a filing cabinet. On your own hardware the ranking inverts: the question is which strong open-weight model actually loads, and how much context it holds once it does.
This guide sorts open-weight models used for coding into four hardware tiers, from each maker's own model card.
What is the best local LLM for coding in 2026?
The best local LLM for coding in 2026 is the largest open-weight model your GPU memory can hold at the context length you actually work at. On a 24 GB card that is Devstral Small 2 or a 4-bit Qwen3-Coder-30B-A3B; on an 80 GB GPU it is gpt-oss-120b or a 4-bit Qwen3-Coder-Next.
How a local LLM for coding differs from a cloud assistant
A local LLM for coding runs inference on hardware you control rather than on a vendor's. You trade some frontier-model quality for that and take on the serving work: picking a quantization, sizing the context cache, and tracking new weights as they ship.
Which local coding model fits your GPU?
Size the weights first. Hugging Face's optimization guide gives the rule: "Loading the weights of a model having X billion parameters requires roughly 2 * X GB of VRAM in bfloat16/float16 precision". Four-bit weights need at least a quarter. Context cache sits on top. Mixture-of-experts models complicate it: Qwen3-Coder-Next activates 3B parameters per token but holds 80B in memory, so size by total parameters.
| Model (maker) | Total / active parameters | Context window | License | Memory (card-stated or calculated) | What the card leads with |
|---|---|---|---|---|---|
| gpt-oss-20b (OpenAI) | 21B / 3.6B | Not on the card | Apache 2.0 | Within 16 GB (card-stated) | Function calling, agentic operations |
| Devstral Small 2 24B (Mistral AI) | 24B | 256K tokens | Apache 2.0 | One RTX 4090 or a 32 GB Mac (card-stated) | SWE-bench Verified 68.0%, per its card |
| Qwen3-Coder-30B-A3B (Qwen) | 30.5B / 3.3B | 262,144, 1M with YaRN | Apache 2.0 | About 61 GB at 16-bit, 15 GB at 4-bit (calculated) | Tool calling; Qwen Code, Cline |
| Qwen3.8-27B (Qwen) | 27B dense | 262,144, 1M extensible | Apache 2.0 | About 54 GB at 16-bit, 14 GB at 4-bit (calculated) | Coding and long-horizon agentic tasks |
| Qwen3-Coder-Next (Qwen) | 80B / 3B | 262,144 tokens | Apache 2.0 | About 160 GB at 16-bit, 40 GB at 4-bit (calculated) | SWE-bench Verified 70.6, per its card |
| gpt-oss-120b (OpenAI) | 117B / 5.1B | Not on the card | Apache 2.0 | One 80 GB GPU (card-stated) | Agentic operations on one GPU |
| Devstral 2 123B (Mistral AI) | 123B | 256K tokens | Modified MIT | About 246 GB at 16-bit (calculated) | Larger sibling of Devstral Small 2 |
| DeepSeek-V4-Flash (DeepSeek) | 284B / 13B | 1M tokens | MIT | Multi-GPU server (calculated) | One-million-token context; three reasoning modes |
| GLM-5.3-Flash (Z.ai) | 320B / 18B | Not on the card | MIT | Multi-GPU server (calculated) | Coding and agentic work |
Capabilities, sizes and licenses as documented by each vendor on 1 October 2026; links in the text. "Card-stated" memory is the maker's own figure. "Calculated" memory is our arithmetic from the parameter count using the Hugging Face rule above, weights only, and is not a vendor figure.

The best free LLM for coding is the one with a permissive license
Every model above is free to download, but "free" is a license question, not a price. Apache 2.0 and MIT carry no user cap and no revenue threshold, which covers Qwen3-Coder, gpt-oss, Devstral Small 2, DeepSeek-V4-Flash and GLM-5.3-Flash. Devstral 2 123B ships under a modified MIT license whose file withdraws the rights from companies above a monthly revenue threshold, so have legal read it before you standardize.
Four hardware tiers, in plain terms
On 16 GB, gpt-oss-20b is the floor: OpenAI's card says it runs within 16 GB of memory. At 24 to 32 GB, Devstral Small 2 24B is the strongest fit, "light enough to run on a single RTX 4090 or a Mac with 32GB RAM" in Mistral's words, at 68.0% on SWE-bench Verified.
On one large GPU, Qwen3-Coder-Next needs about 40 GB for weights at 4-bit by the same arithmetic, and OpenAI says gpt-oss-120b fits a single 80 GB GPU. DeepSeek-V4-Flash and GLM-5.3-Flash are the fourth tier: a multi-GPU server.
Which local models are good enough for agentic coding?
The best local LLM for agentic coding needs three things beyond completion quality: reliable tool calling, a context window large enough for a repository map plus several files, and a scaffold the model was trained against.
Among local LLM models for coding, Qwen3-Coder-Next is the clearest agentic pick for a single large GPU. Its card says it was "designed specifically for coding agents and local development", names the scaffolds it fits, including Claude Code, Qwen Code, Kilo, Trae and Cline, and reports 70.6 on SWE-bench Verified with 36.2 on Terminal-Bench 2.0.
Qwen3-Coder-30B-A3B is the same idea one tier down: its card says it "excels in tool calling capabilities" and supports Qwen Code and Cline through a purpose-built function-call format. Devstral Small 2 lists Mistral Vibe CLI, Cline, Kilo Code, Claude Code, OpenHands and SWE Agent. These scores are vendor-reported on the model cards, so treat them as a shortlist signal, not a result.
Context length is the quiet constraint
An agent that reads ten files and a test log burns context quickly. Qwen's cards say what to do: "If you encounter out-of-memory (OOM) issues, consider reducing the context length to a shorter value, such as 32,768."
Ollama, LM Studio or vLLM: how should you run a local coding model?
Match the runtime to the number of people using the model. One developer on a laptop wants a desktop runner; a team sharing a GPU server wants a serving engine with batching. All four below expose an OpenAI-compatible endpoint, so your editor configuration barely changes.
- Ollama supports "a subset of the OpenAI API" and serves on
http://localhost:11434/v1/. Our hands-on guide to running an LLM locally with Ollama walks through the first model. - LM Studio starts a local server with
lms server startand exposes "chat, responses, embeddings, and other familiar OpenAI-style endpoints". - llama.cpp is the C/C++ engine underneath much of this category, and its repository states that "Apple silicon is a first-class citizen".
- vLLM is what you want once several engineers share one model, because its documentation lists "continuous batching of incoming requests". Our comparison of vLLM and Ollama for enterprise serving covers that switch.
The best free local LLM for coding is therefore a pair, not one choice: an Apache 2.0 or MIT model plus a runtime that fits the audience.
Is a local LLM good enough to replace a cloud coding assistant?
For most day-to-day work on the right hardware, yes; for the hardest multi-file refactors, test before you assume. The Devstral Small 2 and Qwen3-Coder-Next cards report 68.0% and 70.6% on SWE-bench Verified, but those are vendor-reported numbers on public tasks, not on your codebase.
The honest comparison is with the tools your team already uses. Cursor describes itself as "a coding agent for building ambitious software" and Anthropic describes Claude Code as "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools". By default, both run inference on hosted models. Choose one of them when nothing in your security review blocks that and you want a managed agent with no serving work; a fuller side-by-side lives in our Cursor vs Claude Code comparison.
You can run both behind one internally hosted gateway.
How do you benchmark a local coding model on your own repo?
Build a fixed task set from your own repository, run two or three candidates against it on the same serving stack, and score them on the same rubric. The best local LLM for coding on a public table can still fail on your framework mix, your internal APIs and your house style.
- Collect 30 to 50 real tasks from merged pull requests: a bug fix, a test to write, a refactor, a migration, a tricky review comment.
- Record the accepted answer for each, so scoring is a comparison and not an opinion.
- Serve every candidate as you would in production, at the same quantization and context length.
- Score four things per task: correctness, whether it compiled or passed tests, tokens consumed, time to first token.
- Measure peak GPU memory under your real concurrency, not one request at a time.
- Re-run the same set when a new model ships. The set is the asset, not the result.
How do teams share one self-hosted coding model safely?
Put the model behind one internally hosted gateway and give every tool a key through it. A model served straight from a workstation has no quota, no audit trail and no way to revoke access when someone leaves.
Four controls do most of the work: per-team authentication and role-based access, token quotas so one agent loop cannot exhaust the GPU, local logging of every request and response, and secret or PII redaction before content reaches the model layer. If some traffic still goes to a hosted model, the gateway is where you draw that line, per team and per repository. Our guide to the best open-source LLMs to self-host covers the wider model catalogue behind that gateway.
What mistakes should you avoid when choosing a local coding model?
Most failed attempts to pick the best local LLM for coding trace to sizing and licensing:
- Sizing memory for the weights only. A 262,144-token context window needs far more cache than the weights table suggests.
- Reading the family name instead of the license file. Mistral ships Devstral Small 2 under Apache 2.0 and Devstral 2 123B under a modified MIT license with a revenue condition.
- Counting active parameters on a mixture-of-experts model. Qwen3-Coder-Next activates 3B per token and still needs all 80B resident.
- Picking on SWE-bench alone. Agentic work also depends on the scaffold.
- Skipping the task set. Without one you cannot tell whether the next release is better for you.
How Origins AI serves local coding models to a whole team
Origins AI (originshq.com) is a US-based AI-augmented engineering company that deploys self-hosted enterprise AI inside the customer's own infrastructure. The Origins AI Coding Tool is described on its product page as a self-hosted AI coding assistant and LLM gateway for your own infrastructure.
According to that page, the gateway proxies every LLM request from engineering tools through one self-hosted endpoint, with routing, rate limiting, cost tracking and access control across every model a team uses, exposed as an OpenAI-compatible REST API. It lists per-team and per-engineer quotas, local logging of every request and response, and automatic redaction of secrets, API keys and PII before content reaches the model layer. Supported models include Meta Llama, Mistral, CodeLlama, DeepSeek Coder or your own.
Deployment modes on the page are on-premise, private cloud in your own AWS, Azure or GCP account, air-gapped and hybrid. In on-premise and air-gapped modes the page states that no source code is sent to any external service; in hybrid mode, the specific code context submitted to the model leaves your network. Teams that want a model tuned on their own code can go further with Origins AI Domain-Specific LLMs, described as fine-tuning behind your firewall in four stages.
Choose the open-source stack above instead when your platform team already runs GPU servers and is comfortable operating vLLM and a gateway in-house.
Talk to an engineer
Tell us your GPU inventory and your language mix, and we will come back with the best local LLM for coding on that hardware and a serving plan. Book a call with our engineering team to walk through it.


