7 models

Commercially Free Local AI Models (Apache 2.0)

Every model here ships under the Apache 2.0 license, so you can run, modify and ship it in a commercial product with no usage restrictions.

LightMost popular

Qwen3 8B

Alibaba · Apache 2.0

One of the most-downloaded mid-size AI models on HuggingFace in 2026, with tens of millions of downloads. Qwen3 8B outperforms the previous Qwen2.5 14B on most benchmarks, supports hybrid thinking mode and tool calling, and runs on any GPU with 6 GB VRAM. Apache 2.0 license — fully commercial.

Min VRAM

5 GB

Context

128K

codingchatreasoningtool use

Run with Ollama

ollama run qwen3:8b
Light

Mistral 7B v0.3

Mistral AI · Apache 2.0

The model that proved 7B parameters can match much larger models on reasoning tasks when trained carefully. Apache 2.0 licensed, so it is fully free to use commercially with no restrictions.

Min VRAM

4 GB

Context

32K

chatanalysissummarization

Run with Ollama

ollama run mistral:7b
Heavy

Mixtral 8x7B

Mistral AI · Apache 2.0

A sparse mixture-of-experts model that activates only 12B parameters per token while carrying 47B total, delivering 70B-class results at a fraction of the compute cost. One of the best open-weight models available under a fully permissive license.

Min VRAM

26 GB

Context

32K

codinganalysischatcreative writing

Run with Ollama

ollama run mixtral:8x7b
Standard

Mistral Nemo 12B

Mistral AI · Apache 2.0

Built in partnership with NVIDIA and trained on a broad multilingual corpus, Nemo is Mistral's best model for non-English tasks across European and Asian languages. The 128K context window makes it practical for processing long documents locally.

Min VRAM

7 GB

Context

128K

chatmultilingualanalysis

Run with Ollama

ollama run mistral-nemo:12b
StandardBest for fine-tuning

Gemma 4 12B

Google · Apache 2.0

Google's latest Gemma generation doubles the context window to 256K tokens and — a major shift from earlier Gemma releases — ships under the Apache 2.0 license. The 12B model is one of the most actively fine-tuned coding bases on HuggingFace in 2026, with hundreds of community GGUF variants. Runs on any GPU with 10 GB VRAM.

Min VRAM

8 GB

Context

256K

codingchatlong documentsanalysis

Run with Ollama

ollama run gemma4:12b
StandardHistoric first

OpenAI gpt-oss-20b

OpenAI · Apache 2.0

OpenAI's first commercially-licensed open-weight model — a historic release under Apache 2.0. MoE architecture keeps only 3.6B parameters active at inference, so it generates tokens at nearly 7B speed despite 20B total params. Supports function calling and web browsing. Over 10M downloads across Ollama and HuggingFace. Pull it with ollama run gpt-oss:20b.

Min VRAM

12 GB

Context

128K

chatcodingreasoningtool use

Run with Ollama

ollama run gpt-oss:20b
Light

Qwen 2.5 Coder 7B

Alibaba · Apache 2.0

A coding-specialist model fine-tuned on a massive corpus of source code across 40+ programming languages, delivering autocomplete and generation quality that rivals dedicated IDE tools. At 7B it is fast enough for real-time code suggestions on a single consumer GPU.

Min VRAM

4 GB

Context

128K

codingcode completiondebugging

Run with Ollama

ollama run qwen2.5-coder:7b

Want to mix in more filters?

Open the faceted model browser to combine VRAM, family, license, developer and task filters, then sort the results your way.

Related collections