Contact Us

Best Local LLM for Coding, Ranked by Hardware (2026)

Oct 3, 202612 min read
Three glass pedestals of rising size holding glowing compute modules: Best Local LLM for Coding, Ranked by Hardware (2026)
best local llm for coding local llm for coding best llm for coding best free llm for coding local llm models for coding best free local llm for coding

TL;DR

  • Hardware decides which model is "best" for you.
  • Agentic coding needs larger context and tool use, not just code completion.
  • Benchmark on your own repository before standardizing.

Last updated: 3 October 2026

Quick Answer: The best local LLM for coding is decided by VRAM: gpt-oss-20b at 16 GB, Devstral Small 2 24B at 24 to 32 GB. One 80 GB GPU opens gpt-oss-120b or a 4-bit Qwen3-Coder-Next, and DeepSeek-V4-Flash and GLM-5.3-Flash need a multi-GPU server. Take the largest model that fits, then test it on your own repository.

Most rankings of the best LLM for coding assume you can rent a GPU the size of a filing cabinet. On your own hardware the ranking inverts: the question is which strong open-weight model actually loads, and how much context it holds once it does.

This guide sorts open-weight models used for coding into four hardware tiers, from each maker's own model card.

What is the best local LLM for coding in 2026?

The best local LLM for coding in 2026 is the largest open-weight model your GPU memory can hold at the context length you actually work at. On a 24 GB card that is Devstral Small 2 or a 4-bit Qwen3-Coder-30B-A3B; on an 80 GB GPU it is gpt-oss-120b or a 4-bit Qwen3-Coder-Next.

How a local LLM for coding differs from a cloud assistant

A local LLM for coding runs inference on hardware you control rather than on a vendor's. You trade some frontier-model quality for that and take on the serving work: picking a quantization, sizing the context cache, and tracking new weights as they ship.

Which local coding model fits your GPU?

Size the weights first. Hugging Face's optimization guide gives the rule: "Loading the weights of a model having X billion parameters requires roughly 2 * X GB of VRAM in bfloat16/float16 precision". Four-bit weights need at least a quarter. Context cache sits on top. Mixture-of-experts models complicate it: Qwen3-Coder-Next activates 3B parameters per token but holds 80B in memory, so size by total parameters.

Model (maker) Total / active parameters Context window License Memory (card-stated or calculated) What the card leads with
gpt-oss-20b (OpenAI) 21B / 3.6B Not on the card Apache 2.0 Within 16 GB (card-stated) Function calling, agentic operations
Devstral Small 2 24B (Mistral AI) 24B 256K tokens Apache 2.0 One RTX 4090 or a 32 GB Mac (card-stated) SWE-bench Verified 68.0%, per its card
Qwen3-Coder-30B-A3B (Qwen) 30.5B / 3.3B 262,144, 1M with YaRN Apache 2.0 About 61 GB at 16-bit, 15 GB at 4-bit (calculated) Tool calling; Qwen Code, Cline
Qwen3.8-27B (Qwen) 27B dense 262,144, 1M extensible Apache 2.0 About 54 GB at 16-bit, 14 GB at 4-bit (calculated) Coding and long-horizon agentic tasks
Qwen3-Coder-Next (Qwen) 80B / 3B 262,144 tokens Apache 2.0 About 160 GB at 16-bit, 40 GB at 4-bit (calculated) SWE-bench Verified 70.6, per its card
gpt-oss-120b (OpenAI) 117B / 5.1B Not on the card Apache 2.0 One 80 GB GPU (card-stated) Agentic operations on one GPU
Devstral 2 123B (Mistral AI) 123B 256K tokens Modified MIT About 246 GB at 16-bit (calculated) Larger sibling of Devstral Small 2
DeepSeek-V4-Flash (DeepSeek) 284B / 13B 1M tokens MIT Multi-GPU server (calculated) One-million-token context; three reasoning modes
GLM-5.3-Flash (Z.ai) 320B / 18B Not on the card MIT Multi-GPU server (calculated) Coding and agentic work

Capabilities, sizes and licenses as documented by each vendor on 1 October 2026; links in the text. "Card-stated" memory is the maker's own figure. "Calculated" memory is our arithmetic from the parameter count using the Hugging Face rule above, weights only, and is not a vendor figure.

Four hardware tiers for local coding models, 16 GB to multi-GPU server, with card-stated and calculated memory

The best free LLM for coding is the one with a permissive license

Every model above is free to download, but "free" is a license question, not a price. Apache 2.0 and MIT carry no user cap and no revenue threshold, which covers Qwen3-Coder, gpt-oss, Devstral Small 2, DeepSeek-V4-Flash and GLM-5.3-Flash. Devstral 2 123B ships under a modified MIT license whose file withdraws the rights from companies above a monthly revenue threshold, so have legal read it before you standardize.

Four hardware tiers, in plain terms

On 16 GB, gpt-oss-20b is the floor: OpenAI's card says it runs within 16 GB of memory. At 24 to 32 GB, Devstral Small 2 24B is the strongest fit, "light enough to run on a single RTX 4090 or a Mac with 32GB RAM" in Mistral's words, at 68.0% on SWE-bench Verified.

On one large GPU, Qwen3-Coder-Next needs about 40 GB for weights at 4-bit by the same arithmetic, and OpenAI says gpt-oss-120b fits a single 80 GB GPU. DeepSeek-V4-Flash and GLM-5.3-Flash are the fourth tier: a multi-GPU server.

Which local models are good enough for agentic coding?

The best local LLM for agentic coding needs three things beyond completion quality: reliable tool calling, a context window large enough for a repository map plus several files, and a scaffold the model was trained against.

Among local LLM models for coding, Qwen3-Coder-Next is the clearest agentic pick for a single large GPU. Its card says it was "designed specifically for coding agents and local development", names the scaffolds it fits, including Claude Code, Qwen Code, Kilo, Trae and Cline, and reports 70.6 on SWE-bench Verified with 36.2 on Terminal-Bench 2.0.

Qwen3-Coder-30B-A3B is the same idea one tier down: its card says it "excels in tool calling capabilities" and supports Qwen Code and Cline through a purpose-built function-call format. Devstral Small 2 lists Mistral Vibe CLI, Cline, Kilo Code, Claude Code, OpenHands and SWE Agent. These scores are vendor-reported on the model cards, so treat them as a shortlist signal, not a result.

Context length is the quiet constraint

An agent that reads ten files and a test log burns context quickly. Qwen's cards say what to do: "If you encounter out-of-memory (OOM) issues, consider reducing the context length to a shorter value, such as 32,768."

Ollama, LM Studio or vLLM: how should you run a local coding model?

Match the runtime to the number of people using the model. One developer on a laptop wants a desktop runner; a team sharing a GPU server wants a serving engine with batching. All four below expose an OpenAI-compatible endpoint, so your editor configuration barely changes.

The best free local LLM for coding is therefore a pair, not one choice: an Apache 2.0 or MIT model plus a runtime that fits the audience.

Is a local LLM good enough to replace a cloud coding assistant?

For most day-to-day work on the right hardware, yes; for the hardest multi-file refactors, test before you assume. The Devstral Small 2 and Qwen3-Coder-Next cards report 68.0% and 70.6% on SWE-bench Verified, but those are vendor-reported numbers on public tasks, not on your codebase.

The honest comparison is with the tools your team already uses. Cursor describes itself as "a coding agent for building ambitious software" and Anthropic describes Claude Code as "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools". By default, both run inference on hosted models. Choose one of them when nothing in your security review blocks that and you want a managed agent with no serving work; a fuller side-by-side lives in our Cursor vs Claude Code comparison.

You can run both behind one internally hosted gateway.

How do you benchmark a local coding model on your own repo?

Build a fixed task set from your own repository, run two or three candidates against it on the same serving stack, and score them on the same rubric. The best local LLM for coding on a public table can still fail on your framework mix, your internal APIs and your house style.

  1. Collect 30 to 50 real tasks from merged pull requests: a bug fix, a test to write, a refactor, a migration, a tricky review comment.
  2. Record the accepted answer for each, so scoring is a comparison and not an opinion.
  3. Serve every candidate as you would in production, at the same quantization and context length.
  4. Score four things per task: correctness, whether it compiled or passed tests, tokens consumed, time to first token.
  5. Measure peak GPU memory under your real concurrency, not one request at a time.
  6. Re-run the same set when a new model ships. The set is the asset, not the result.

How do teams share one self-hosted coding model safely?

Put the model behind one internally hosted gateway and give every tool a key through it. A model served straight from a workstation has no quota, no audit trail and no way to revoke access when someone leaves.

Four controls do most of the work: per-team authentication and role-based access, token quotas so one agent loop cannot exhaust the GPU, local logging of every request and response, and secret or PII redaction before content reaches the model layer. If some traffic still goes to a hosted model, the gateway is where you draw that line, per team and per repository. Our guide to the best open-source LLMs to self-host covers the wider model catalogue behind that gateway.

What mistakes should you avoid when choosing a local coding model?

Most failed attempts to pick the best local LLM for coding trace to sizing and licensing:

How Origins AI serves local coding models to a whole team

Origins AI (originshq.com) is a US-based AI-augmented engineering company that deploys self-hosted enterprise AI inside the customer's own infrastructure. The Origins AI Coding Tool is described on its product page as a self-hosted AI coding assistant and LLM gateway for your own infrastructure.

According to that page, the gateway proxies every LLM request from engineering tools through one self-hosted endpoint, with routing, rate limiting, cost tracking and access control across every model a team uses, exposed as an OpenAI-compatible REST API. It lists per-team and per-engineer quotas, local logging of every request and response, and automatic redaction of secrets, API keys and PII before content reaches the model layer. Supported models include Meta Llama, Mistral, CodeLlama, DeepSeek Coder or your own.

Deployment modes on the page are on-premise, private cloud in your own AWS, Azure or GCP account, air-gapped and hybrid. In on-premise and air-gapped modes the page states that no source code is sent to any external service; in hybrid mode, the specific code context submitted to the model leaves your network. Teams that want a model tuned on their own code can go further with Origins AI Domain-Specific LLMs, described as fine-tuning behind your firewall in four stages.

Choose the open-source stack above instead when your platform team already runs GPU servers and is comfortable operating vLLM and a gateway in-house.

Talk to an engineer

Tell us your GPU inventory and your language mix, and we will come back with the best local LLM for coding on that hardware and a serving plan. Book a call with our engineering team to walk through it.

Frequently Asked Questions

Can I run a coding LLM locally?
Yes, on most recent hardware. A 16 GB GPU or Mac runs gpt-oss-20b, which OpenAI's model card says fits within 16 GB of memory, and a 32 GB machine runs Devstral Small 2 24B. The table above starts at the 16 GB tier.
What should I run on a Mac with Apple silicon?
On a 32 GB Mac, Devstral Small 2 24B is the strongest documented fit: Mistral's card names a Mac with 32 GB RAM explicitly. With 64 GB or more, a 4-bit Qwen3-Coder-Next is worth testing: its weights calculate to about 40 GB. Apple silicon does well because llama.cpp treats it as a first-class Metal target.
How much VRAM do I need for a local coding model?
Start from the Hugging Face rule of roughly 2 GB per billion parameters at 16-bit; 4-bit weights need at least a quarter of that. A 24B model needs about 48 GB at 16-bit, or around 12 GB at 4-bit, before context. Add headroom for the cache, which grows with context length and concurrent users.
Are local coding LLMs free to use commercially?
Most are. Qwen3-Coder, gpt-oss, Devstral Small 2, DeepSeek-V4-Flash and GLM-5.3-Flash ship under Apache 2.0 or MIT, which carry no user or revenue cap. Devstral 2 123B uses a modified MIT license. Read the license file in the exact repository you deploy from, because one family can ship several licenses, and check whether fine-tuned derivatives carry naming or notice duties.
Which LLM is best for coding, local or cloud?
Decide by constraint, not by score. If a security review blocks sending code to a vendor, a local model is the only option and the question becomes which one fits your GPUs. If nothing blocks it, test a hosted agent such as Claude Code against your local pick on your hardest multi-file tasks. You can also route both and split per repository.
Book a call

About the Author

Apoorva Kumar is Co-Founder and CEO of Origins AI (originshq.com), an AI engineering partner for product teams building AI workflows, AI agents and LLM integrations. A CSE graduate of IIT Kharagpur, Apoorva previously built and scaled technology at Sony, NuCash, YesMadam and FrontPage.