Contact Us

Best Open-Source LLMs to Self-Host for Business Use (2026)

Sep 25, 202610 min read
Origins AI article banner with the title: Best Open-Source LLMs to Self-Host for Business Use (2026)
best open source llm best open source llm for coding best open source llm to run locally open-weight model llm gpu requirements

TL;DR

  • Read the license file, not the family name, because Qwen and GLM ship some releases under Apache 2.0 or MIT and others under custom terms.
  • Weights need roughly 2 GB of GPU memory per billion parameters at 16-bit, and mixture-of-experts models must be sized by total parameters.
  • Run two or three shortlisted models on 50 to 200 real questions from your own work, served the same way as in production.

Quick Answer: The best open-source LLMs to self-host for business in 2026 come from seven families: Qwen, DeepSeek, GLM, Mistral, Gemma, gpt-oss and Llama. Shortlist by license first, since Apache 2.0 and MIT carry no user caps, then by task and GPU memory. gpt-oss-120b fits one 80 GB GPU; DeepSeek V4 needs a multi-GPU server.

Most "best model" lists rank by benchmark score. A business running weights on its own hardware has three harder filters: what the license allows at your scale, how well the model handles your workload, and how many GPUs it needs at your context length.

This guide sorts the leading open-weight model families by those three filters, using each maker's own model card. It's written for the CTO, platform lead and security reviewer who sign off together.

Which open-source LLMs are best for business use right now?

The best open-source LLM for most businesses in 2026 is a permissively licensed mid-size model: Qwen3.8-27B, Mistral Small 4, Gemma 4 or gpt-oss-120b. Pick DeepSeek V4 or GLM-5.3-Flash when you need frontier-level reasoning and can run a multi-GPU server.

Here is the shortlist by use case:

Model (maker) Total / active parameters Context window License Permissive license Strongest fit (per the model card)
Qwen3.8-27B (Qwen) 27B dense 262K tokens Apache 2.0 Yes Coding, agents, image and video input
Qwen3.8-Flash-Next (Qwen) 125B / 6B, plus 51B n-gram embeddings 262K tokens Qwen Community License No Efficient long-context serving
DeepSeek-V4-Pro (DeepSeek) 1.6T / 49B 1M tokens MIT Yes Maximum reasoning and coding
DeepSeek-V4-Flash (DeepSeek) 284B / 13B 1M tokens MIT Yes Long context at lower compute
GLM-5.3-Flash (Z.ai) 320B / 18B 1M tokens MIT Yes Coding and agentic work
Mistral Small 4 (Mistral AI) 119B / 6.5B 256K tokens Apache 2.0 Yes Chat, coding, documents, function calling
Gemma 4 31B (Google DeepMind) 31B dense 256K tokens Apache 2.0 Yes Workstation deployment; smaller sizes for edge
gpt-oss-120b (OpenAI) 117B / 5.1B 131K tokens Apache 2.0 Yes Reasoning on one 80 GB GPU
Llama 4 Scout / Maverick (Meta) 109B or 400B / 17B 10M / 1M tokens Llama 4 Community License No Multilingual text and image, very long context

Licenses, sizes and context lengths as documented by each vendor on 25 September 2026, read from each model card on Hugging Face; links in the text.

How do licenses differ across open-source LLMs?

Open-weight licenses fall into three groups: permissive (Apache 2.0 or MIT, no user or revenue caps), community licenses with scale thresholds, and revenue-capped licenses. Legal should read the license file in the repository, because one family can ship under several licenses.

Permissive: Apache 2.0 and MIT

Qwen3.5, Qwen3.8-27B, Mistral Small 4, Gemma 4 and gpt-oss ship under Apache 2.0; DeepSeek V4 and GLM-5.3-Flash under MIT. Both let you use, modify, fine-tune and sell products built on the weights, with notice duties but no cap on users or revenue.

Community licenses with thresholds

The Llama 4 Community License requires a separate license from Meta if your products had more than 700 million monthly active users on the release date. It also requires a "Built with Llama" notice, and any fine-tuned model you distribute must start its name with "Llama".

Qwen's newer large releases use their own terms. The Qwen Community License asks for the model name on your interface above 100 million monthly active users or a stated revenue level. It also requires a separate license if you run a model-as-a-service or AI work-assistant business, while internal use stays exempt.

Z.ai's GLM-5.3 license (the full model, not the MIT-licensed Flash) requires model-as-a-service operators above a revenue threshold to pass a Z.ai security review first.

Revenue-capped licenses

Mistral Medium 3.5 uses a modified MIT license that withdraws all rights from companies above a monthly revenue level. Larger companies need a commercial agreement with Mistral AI.

What legal should check in every license file:

Which open-source models fit coding, chat and document tasks?

Task fit matters more than a leaderboard rank. Judge each model on four things: output quality on your own prompts, the GPU memory it needs, its license, and how it behaves in production (tool calling, structured output, refusal rate). Public rankings such as LMArena's leaderboard help you build a shortlist, not make the decision.

Coding and agents

For the best open-source LLM for coding, start with GLM-5.3-Flash, DeepSeek V4 and Qwen3.8-27B, whose cards lead with vendor-reported coding results. Mistral's Devstral Small 2 (24B, Apache 2.0) is a smaller option.

Chat and internal assistants

Mistral Small 4 and Qwen3.8-27B switch between a fast reply mode and a reasoning mode per request, which keeps simple answers quick. gpt-oss-120b offers three reasoning levels on one GPU.

Documents and retrieval

Long context helps, but retrieval usually beats pasting whole archives into a prompt. DeepSeek V4 and GLM-5.3-Flash accept one million tokens and Llama 4 Scout 10 million, yet cache memory grows with every token.

Choose Llama 4 when your stack is already built on Llama tooling and you sit well under the user threshold. Choose an Apache 2.0 or MIT model when you plan to host the model for customers or redistribute a fine-tune.

Which GPUs do you need to self-host these models?

LLM GPU requirements start with one rule: weights need roughly 2 GB of GPU memory per billion parameters at 16-bit precision, and about a quarter of that at 4-bit. Add headroom for the context cache, which grows with context length and concurrent users.

The 16-bit rule comes from Hugging Face's guide to LLM memory requirements. Mixture-of-experts (MoE) models activate only a few billion parameters per token, but every expert must still sit in memory, so size by total parameters.

Model size (total parameters) Weights at 16-bit Weights at 4-bit Typical GPU setup
9B to 12B (Qwen3.5-9B, Gemma 4 12B) 19 to 24 GB 5 to 6 GB One workstation GPU
21B to 31B (gpt-oss-20b, Qwen3.8-27B, Gemma 4 31B) 42 to 63 GB 11 to 16 GB One 80 GB card; a 24 GB card at 4-bit
70B dense About 140 GB About 35 GB Two 80 GB GPUs at 16-bit; one 48 GB card at 4-bit
109B to 119B MoE (Llama 4 Scout, gpt-oss-120b, Mistral Small 4) 218 to 238 GB 55 to 60 GB One 80 GB or 96 GB GPU at 4-bit
284B to 400B MoE (DeepSeek-V4-Flash, GLM-5.3-Flash, Llama 4 Maverick) 568 to 800 GB 142 to 200 GB A multi-GPU server
1.6T MoE (DeepSeek-V4-Pro) About 3.2 TB About 800 GB An 8-GPU server with 141 GB cards at 4-bit; several nodes at 16-bit

Weights only, from the rule above; context cache and serving overhead come on top.

Data-center and workstation GPU classes

Data-center cards carry the most memory per GPU: NVIDIA lists 80 GB on the H100 SXM and 141 GB on the H200. The L40S offers 48 GB, and the RTX PRO 6000 Blackwell workstation card offers 96 GB. OpenAI says gpt-oss-120b fits a single 80 GB GPU and gpt-oss-20b runs within 16 GB, because both ship pre-quantized.

Quantizing to 8-bit or 4-bit fits a model onto fewer cards; see our walkthrough on quantizing models with GGUF.

If several teams will call two or three models, serve them behind one internally hosted gateway so routing, quotas and logs live in one place. Our comparison of self-hosted LLM gateways covers the options. The engine behind that gateway is its own decision, and how vLLM and Ollama compare for enterprise serving lays out the tradeoffs.

How do you test an open-source model on your own data?

Build a private evaluation set from your own work, run two or three shortlisted models against it on the same serving stack, and score accuracy, latency and memory. A model that tops a public ranking can still miss your terminology and your document formats.

  1. Collect 50 to 200 real questions from tickets, documents or chat logs, with the answer an expert would accept.
  2. Add hard cases on purpose: ambiguous questions, long documents and requests the model should refuse.
  3. Serve each candidate the same way you'll run it in production, at the same quantization and context length.
  4. Measure answer accuracy, time to first token, throughput under your expected concurrency and peak GPU memory.
  5. Send the license file to legal while the test runs, so the winner isn't blocked a month later.

If you're weighing whether to fine-tune a domain model yourself or with a development firm, test the base model on your own data first. The evaluation set shows whether retrieval closes the gap or fine-tuning is worth it, which matters in fintech and healthcare, where wording is precise.

When is a smaller open model the better choice?

A smaller open model is the better choice when the task is narrow, latency matters more than breadth, or the model must run on one GPU or on a device. Classification, extraction, routing and short summaries rarely need a 400B model.

Small models are also cheaper to fine-tune and faster to re-test when your data changes. Gemma 4's E2B and E4B are built for laptops and mobile devices, and Qwen3.5 goes down to 0.8B.

Choose a small model when you can describe the task in one sentence and measure it. Choose a large model for open-ended questions across many domains.

What mistakes should you avoid when picking an open-source LLM?

The costliest mistakes come from treating "open" as "no obligations" and sizing hardware for the weights alone:

Where you run the model matters too; see our guide to on-premise, private cloud and air-gapped AI.

How Origins AI deploys open models inside your environment

Origins AI (originshq.com) is a US-based AI-augmented engineering company that deploys self-hosted enterprise AI inside the customer's infrastructure. The Origins AI Domain-Specific LLMs offering trains or fine-tunes models on your proprietary data behind your firewall.

The product page describes four steps: audit and scope, infrastructure setup in your cloud or on-premise, model training (fine-tuning or training from scratch), then integration and iteration. Every product in the Origins AI catalog supports on-premise or private cloud deployment, plus air-gapped mode, with the customer choosing the model.

In on-premise and air-gapped modes, no data leaves your network; hybrid mode sends the submitted context to a hosted model. The site lists encryption at rest and in transit, secure authentication, continuous security monitoring and least-privilege access.

The Origins AI Coding Tool adds a self-hosted gateway with an OpenAI-compatible API that routes to Llama, Mistral, DeepSeek Coder or your own model. Origins AI reports a 40 to 60 percent reduction in knowledge work time as a typical return.

If your team already runs GPU clusters and has ML engineers to serve and evaluate models, running an Apache 2.0 model on your own serving stack is the better fit.

Talk to an engineer

If you're shortlisting open-weight models for a private deployment, book a call with our engineering team to walk through licensing, hardware sizing and an evaluation plan for your data.

Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.

Frequently Asked Questions

How much GPU memory does a 70B model need?
A 70B dense model needs about 140 GB of GPU memory for its weights at 16-bit precision, so plan on two 80 GB GPUs. At 4-bit it drops to about 35 GB, which fits one 48 GB card, plus headroom for the context cache.
What is the best open-source LLM to run locally?
For a single machine, gpt-oss-20b is a strong starting point because OpenAI says it runs within 16 GB of memory. Gemma 4 12B and Qwen3.5-9B are good alternatives on a 16 GB card at 4-bit, and Gemma 4's E2B and E4B are designed for laptops and phones.
Can open-source LLMs be used commercially?
Yes, most can. Apache 2.0 and MIT models such as Mistral Small 4, gpt-oss, Gemma 4 and DeepSeek V4 allow commercial use without user caps. Llama 4 and some Qwen and GLM releases add thresholds, naming duties or separate terms for hosted-model businesses, so read the license file for the exact version you deploy.
Are open-source LLMs as good as ChatGPT?
On many business tasks the gap is now small, and the largest open-weight models publish scores close to frontier closed models. Those numbers are vendor-reported, though. Run the same 100 questions through a hosted model and your open-weight candidate, then compare accuracy and latency side by side.
What open-weight model should a software team try first for code generation?
On a multi-GPU server, GLM-5.3-Flash and DeepSeek V4 lead their model cards with coding and agent results, under MIT licenses. On one GPU, Qwen3.8-27B at 4-bit or Devstral Small 2 (24B) is more practical. Test each on your own repository, because the language and framework mix changes the ranking.
Can I fine-tune an open-source LLM on company data?
Yes. Apache 2.0 and MIT licenses allow fine-tuning, and parameter-efficient methods such as LoRA make it possible on a single GPU for mid-size models. Test retrieval first, because it keeps answers current without retraining, and keep training data and adapters inside your own environment.
Book a call

About the Author

Apoorva Kumar is Co-Founder and CEO of Origins AI (originshq.com), an AI engineering partner for product teams building AI workflows, AI agents and LLM integrations. A CSE graduate of IIT Kharagpur, Apoorva previously built and scaled technology at Sony, NuCash, YesMadam and FrontPage.