Qwen3.8-27B
Alibaba · Qwen
Alibaba's newest Qwen release, published August 14, 2026 under Apache 2.0. A 27B dense model that adds native image and video understanding on top of text, with a 256K context window long enough for hour-scale video or entire codebases. At Q4 it needs about 18 GB of VRAM, so it fits a single 24 GB card with room to spare. Pull it with ollama run qwen3.8:27b.
Params
27B
License
Apache 2.0
Context
256K
Min VRAM
18 GB
Min RAM
32 GB
Tier
Heavy
Run locally with Ollama
ollama run qwen3.8:27bOnce running, Ollama exposes an OpenAI-compatible endpoint at localhost:11434/v1.
Will it run on my hardware?
Qwen3.8-27B needs about 18 GB of fast memory at Q4. Here is what in our hardware database clears that bar.
Runs at full speed · 6
Runs, but slower on shared memory · 3
6 other tracked configurations do not have enough memory.
Hardware links are affiliate links. We earn a small commission if you buy through them — at no extra cost to you. Disclosure
Specifications
| Parameters | 27B |
|---|---|
| Context window | 256K |
| License | Apache 2.0 |
| Min VRAM (Q4) | 18 GB |
| Min RAM | 32 GB |
| Recommended tier | Heavy |
| Best for | codingchatanalysislong documents |
Variants on Ollama
Ollama serves the Q4_K_M quantization by default. Higher-precision tags exist for users with more memory.
Qwen3.8-27B
ollama run qwen3.8:27bollama pull qwen3.8:q8_0ollama pull qwen3.8:fp16Exact quantization tags vary per model — run ollama show qwen3.8 or check the tags page for the full list.
Popularity trend (illustrative)
Illustrative adoption trend, not actual download counts. We do not publish a download number for any model because we have no honest source for it.
Similar models
Mixtral 8x7B
Mistral AI · Apache 2.0
A sparse mixture-of-experts model that activates only 12B parameters per token while carrying 47B total, delivering 70B-class results at a fraction of the compute cost. One of the best open-weight models available under a fully permissive license.
Min VRAM
26 GB
Context
32K
Run with Ollama
ollama run mixtral:8x7bGemma 3 27B
Google · Gemma Terms of Use
The largest Gemma 3 model and one of the best open-weight models in the 20–30B range. Fits in a single RTX 4090 (24 GB) with room to spare at Q4 quantization, and delivers instruction-following quality that rivals models twice its size.
Min VRAM
15 GB
Context
128K
Run with Ollama
ollama run gemma3:27bGemma 4 12B
Google · Apache 2.0
Google's latest Gemma generation doubles the context window to 256K tokens and — a major shift from earlier Gemma releases — ships under the Apache 2.0 license. The 12B model is one of the most actively fine-tuned coding bases on HuggingFace in 2026, with hundreds of community GGUF variants. Runs on any GPU with 10 GB VRAM.
Min VRAM
8 GB
Context
256K
Run with Ollama
ollama run gemma4:12b