Best Qwen Models to Run Locally
Alibaba's Qwen line includes one of the most-downloaded mid-size models of 2026 and a dedicated coding variant, all under the permissive Apache 2.0 license.
Qwen3 8B
Alibaba · Apache 2.0
One of the most-downloaded mid-size AI models on HuggingFace in 2026, with tens of millions of downloads. Qwen3 8B outperforms the previous Qwen2.5 14B on most benchmarks, supports hybrid thinking mode and tool calling, and runs on any GPU with 6 GB VRAM. Apache 2.0 license — fully commercial.
Min VRAM
5 GB
Context
128K
Run with Ollama
ollama run qwen3:8bQwen 2.5 Coder 7B
Alibaba · Apache 2.0
A coding-specialist model fine-tuned on a massive corpus of source code across 40+ programming languages, delivering autocomplete and generation quality that rivals dedicated IDE tools. At 7B it is fast enough for real-time code suggestions on a single consumer GPU.
Min VRAM
4 GB
Context
128K
Run with Ollama
ollama run qwen2.5-coder:7bQwen3-TTS 1.7B CustomVoice
Alibaba · Apache 2.0
Alibaba's Qwen3-TTS with 9 premium preset voices and instruction-based control over tone, emotion and prosody across 10 major languages. Streaming generation gets end-to-end latency as low as 97ms. Apache 2.0 licensed — fully free for commercial use, unlike most other speech models on this page.
Min VRAM
Not published — FlashAttention recommended to reduce memory use
Speed
End-to-end synthesis latency as low as 97ms (streaming)
Install & run
pip install -U qwen-ttsWant to mix in more filters?
Open the faceted model browser to combine VRAM, family, license, developer and task filters, then sort the results your way.