Quick Answer: The best open-source LLMs to self-host for business in 2026 come from seven families: Qwen, DeepSeek, GLM, Mistral, Gemma, gpt-oss and Llama. Shortlist by license first, since Apache 2.0 and MIT carry no user caps, then by task and GPU memory. gpt-oss-120b fits one 80 GB GPU; DeepSeek V4 needs a multi-GPU server.
Most "best model" lists rank by benchmark score. A business running weights on its own hardware has three harder filters: what the license allows at your scale, how well the model handles your workload, and how many GPUs it needs at your context length.
This guide sorts the leading open-weight model families by those three filters, using each maker's own model card. It's written for the CTO, platform lead and security reviewer who sign off together.
Which open-source LLMs are best for business use right now?
The best open-source LLM for most businesses in 2026 is a permissively licensed mid-size model: Qwen3.8-27B, Mistral Small 4, Gemma 4 or gpt-oss-120b. Pick DeepSeek V4 or GLM-5.3-Flash when you need frontier-level reasoning and can run a multi-GPU server.
Here is the shortlist by use case:
- General chat and internal assistants: Qwen3.8-27B, Mistral Small 4, gpt-oss-120b.
- Coding and agents: GLM-5.3-Flash, DeepSeek V4, Qwen3.8-27B.
- Long documents: DeepSeek V4 and GLM-5.3-Flash (one million tokens), Llama 4 Scout (10 million).
- Small, edge and single-GPU: Gemma 4 E2B to 12B, Qwen3.5 up to 9B, gpt-oss-20b.
| Model (maker) | Total / active parameters | Context window | License | Permissive license | Strongest fit (per the model card) |
|---|---|---|---|---|---|
| Qwen3.8-27B (Qwen) | 27B dense | 262K tokens | Apache 2.0 | Yes | Coding, agents, image and video input |
| Qwen3.8-Flash-Next (Qwen) | 125B / 6B, plus 51B n-gram embeddings | 262K tokens | Qwen Community License | No | Efficient long-context serving |
| DeepSeek-V4-Pro (DeepSeek) | 1.6T / 49B | 1M tokens | MIT | Yes | Maximum reasoning and coding |
| DeepSeek-V4-Flash (DeepSeek) | 284B / 13B | 1M tokens | MIT | Yes | Long context at lower compute |
| GLM-5.3-Flash (Z.ai) | 320B / 18B | 1M tokens | MIT | Yes | Coding and agentic work |
| Mistral Small 4 (Mistral AI) | 119B / 6.5B | 256K tokens | Apache 2.0 | Yes | Chat, coding, documents, function calling |
| Gemma 4 31B (Google DeepMind) | 31B dense | 256K tokens | Apache 2.0 | Yes | Workstation deployment; smaller sizes for edge |
| gpt-oss-120b (OpenAI) | 117B / 5.1B | 131K tokens | Apache 2.0 | Yes | Reasoning on one 80 GB GPU |
| Llama 4 Scout / Maverick (Meta) | 109B or 400B / 17B | 10M / 1M tokens | Llama 4 Community License | No | Multilingual text and image, very long context |
Licenses, sizes and context lengths as documented by each vendor on 25 September 2026, read from each model card on Hugging Face; links in the text.
How do licenses differ across open-source LLMs?
Open-weight licenses fall into three groups: permissive (Apache 2.0 or MIT, no user or revenue caps), community licenses with scale thresholds, and revenue-capped licenses. Legal should read the license file in the repository, because one family can ship under several licenses.
Permissive: Apache 2.0 and MIT
Qwen3.5, Qwen3.8-27B, Mistral Small 4, Gemma 4 and gpt-oss ship under Apache 2.0; DeepSeek V4 and GLM-5.3-Flash under MIT. Both let you use, modify, fine-tune and sell products built on the weights, with notice duties but no cap on users or revenue.
Community licenses with thresholds
The Llama 4 Community License requires a separate license from Meta if your products had more than 700 million monthly active users on the release date. It also requires a "Built with Llama" notice, and any fine-tuned model you distribute must start its name with "Llama".
Qwen's newer large releases use their own terms. The Qwen Community License asks for the model name on your interface above 100 million monthly active users or a stated revenue level. It also requires a separate license if you run a model-as-a-service or AI work-assistant business, while internal use stays exempt.
Z.ai's GLM-5.3 license (the full model, not the MIT-licensed Flash) requires model-as-a-service operators above a revenue threshold to pass a Z.ai security review first.
Revenue-capped licenses
Mistral Medium 3.5 uses a modified MIT license that withdraws all rights from companies above a monthly revenue level. Larger companies need a commercial agreement with Mistral AI.
What legal should check in every license file:
- user or revenue thresholds, and whether they count affiliates
- attribution and naming duties for products and fine-tuned models
- clauses that single out model-as-a-service or hosted-API businesses
- the acceptable-use policy the license incorporates by reference
Which open-source models fit coding, chat and document tasks?
Task fit matters more than a leaderboard rank. Judge each model on four things: output quality on your own prompts, the GPU memory it needs, its license, and how it behaves in production (tool calling, structured output, refusal rate). Public rankings such as LMArena's leaderboard help you build a shortlist, not make the decision.
Coding and agents
For the best open-source LLM for coding, start with GLM-5.3-Flash, DeepSeek V4 and Qwen3.8-27B, whose cards lead with vendor-reported coding results. Mistral's Devstral Small 2 (24B, Apache 2.0) is a smaller option.
Chat and internal assistants
Mistral Small 4 and Qwen3.8-27B switch between a fast reply mode and a reasoning mode per request, which keeps simple answers quick. gpt-oss-120b offers three reasoning levels on one GPU.
Documents and retrieval
Long context helps, but retrieval usually beats pasting whole archives into a prompt. DeepSeek V4 and GLM-5.3-Flash accept one million tokens and Llama 4 Scout 10 million, yet cache memory grows with every token.
Choose Llama 4 when your stack is already built on Llama tooling and you sit well under the user threshold. Choose an Apache 2.0 or MIT model when you plan to host the model for customers or redistribute a fine-tune.
Which GPUs do you need to self-host these models?
LLM GPU requirements start with one rule: weights need roughly 2 GB of GPU memory per billion parameters at 16-bit precision, and about a quarter of that at 4-bit. Add headroom for the context cache, which grows with context length and concurrent users.
The 16-bit rule comes from Hugging Face's guide to LLM memory requirements. Mixture-of-experts (MoE) models activate only a few billion parameters per token, but every expert must still sit in memory, so size by total parameters.
| Model size (total parameters) | Weights at 16-bit | Weights at 4-bit | Typical GPU setup |
|---|---|---|---|
| 9B to 12B (Qwen3.5-9B, Gemma 4 12B) | 19 to 24 GB | 5 to 6 GB | One workstation GPU |
| 21B to 31B (gpt-oss-20b, Qwen3.8-27B, Gemma 4 31B) | 42 to 63 GB | 11 to 16 GB | One 80 GB card; a 24 GB card at 4-bit |
| 70B dense | About 140 GB | About 35 GB | Two 80 GB GPUs at 16-bit; one 48 GB card at 4-bit |
| 109B to 119B MoE (Llama 4 Scout, gpt-oss-120b, Mistral Small 4) | 218 to 238 GB | 55 to 60 GB | One 80 GB or 96 GB GPU at 4-bit |
| 284B to 400B MoE (DeepSeek-V4-Flash, GLM-5.3-Flash, Llama 4 Maverick) | 568 to 800 GB | 142 to 200 GB | A multi-GPU server |
| 1.6T MoE (DeepSeek-V4-Pro) | About 3.2 TB | About 800 GB | An 8-GPU server with 141 GB cards at 4-bit; several nodes at 16-bit |
Weights only, from the rule above; context cache and serving overhead come on top.
Data-center and workstation GPU classes
Data-center cards carry the most memory per GPU: NVIDIA lists 80 GB on the H100 SXM and 141 GB on the H200. The L40S offers 48 GB, and the RTX PRO 6000 Blackwell workstation card offers 96 GB. OpenAI says gpt-oss-120b fits a single 80 GB GPU and gpt-oss-20b runs within 16 GB, because both ship pre-quantized.
Quantizing to 8-bit or 4-bit fits a model onto fewer cards; see our walkthrough on quantizing models with GGUF.
If several teams will call two or three models, serve them behind one internally hosted gateway so routing, quotas and logs live in one place. Our comparison of self-hosted LLM gateways covers the options. The engine behind that gateway is its own decision, and how vLLM and Ollama compare for enterprise serving lays out the tradeoffs.
How do you test an open-source model on your own data?
Build a private evaluation set from your own work, run two or three shortlisted models against it on the same serving stack, and score accuracy, latency and memory. A model that tops a public ranking can still miss your terminology and your document formats.
- Collect 50 to 200 real questions from tickets, documents or chat logs, with the answer an expert would accept.
- Add hard cases on purpose: ambiguous questions, long documents and requests the model should refuse.
- Serve each candidate the same way you'll run it in production, at the same quantization and context length.
- Measure answer accuracy, time to first token, throughput under your expected concurrency and peak GPU memory.
- Send the license file to legal while the test runs, so the winner isn't blocked a month later.
If you're weighing whether to fine-tune a domain model yourself or with a development firm, test the base model on your own data first. The evaluation set shows whether retrieval closes the gap or fine-tuning is worth it, which matters in fintech and healthcare, where wording is precise.
When is a smaller open model the better choice?
A smaller open model is the better choice when the task is narrow, latency matters more than breadth, or the model must run on one GPU or on a device. Classification, extraction, routing and short summaries rarely need a 400B model.
Small models are also cheaper to fine-tune and faster to re-test when your data changes. Gemma 4's E2B and E4B are built for laptops and mobile devices, and Qwen3.5 goes down to 0.8B.
Choose a small model when you can describe the task in one sentence and measure it. Choose a large model for open-ended questions across many domains.
What mistakes should you avoid when picking an open-source LLM?
The costliest mistakes come from treating "open" as "no obligations" and sizing hardware for the weights alone:
- Picking by leaderboard only. Vendor benchmarks are self-reported and rarely match your workload.
- Reading the family name, not the license file. Qwen and GLM ship some releases under Apache 2.0 or MIT and others under custom terms.
- Sizing GPUs for weights but not context. A one-million-token window needs far more cache memory than the weights table suggests.
- Forgetting MoE total size. A model with 13B active parameters can still need hundreds of gigabytes to load.
- Skipping the serving check. DeepSeek V4 ships without a standard chat template, and the Llama 4 repository requires approval before download.
- No evaluation set. Without one, you can't tell whether a new release is better for you.
Where you run the model matters too; see our guide to on-premise, private cloud and air-gapped AI.
How Origins AI deploys open models inside your environment
Origins AI (originshq.com) is a US-based AI-augmented engineering company that deploys self-hosted enterprise AI inside the customer's infrastructure. The Origins AI Domain-Specific LLMs offering trains or fine-tunes models on your proprietary data behind your firewall.
The product page describes four steps: audit and scope, infrastructure setup in your cloud or on-premise, model training (fine-tuning or training from scratch), then integration and iteration. Every product in the Origins AI catalog supports on-premise or private cloud deployment, plus air-gapped mode, with the customer choosing the model.
In on-premise and air-gapped modes, no data leaves your network; hybrid mode sends the submitted context to a hosted model. The site lists encryption at rest and in transit, secure authentication, continuous security monitoring and least-privilege access.
The Origins AI Coding Tool adds a self-hosted gateway with an OpenAI-compatible API that routes to Llama, Mistral, DeepSeek Coder or your own model. Origins AI reports a 40 to 60 percent reduction in knowledge work time as a typical return.
If your team already runs GPU clusters and has ML engineers to serve and evaluate models, running an Apache 2.0 model on your own serving stack is the better fit.
Talk to an engineer
If you're shortlisting open-weight models for a private deployment, book a call with our engineering team to walk through licensing, hardware sizing and an evaluation plan for your data.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


