Best Gemma Models to Run Locally
Google's Gemma family runs from a 4B model on integrated graphics up to the 27B heavyweight, with some of the longest context windows in the catalog.
Gemma 4 12B
Google · Apache 2.0
Google's latest Gemma generation doubles the context window to 256K tokens and — a major shift from earlier Gemma releases — ships under the Apache 2.0 license. The 12B model is one of the most actively fine-tuned coding bases on HuggingFace in 2026, with hundreds of community GGUF variants. Runs on any GPU with 10 GB VRAM.
Min VRAM
8 GB
Context
256K
Run with Ollama
ollama run gemma4:12bGemma 3 4B
Google · Gemma Terms of Use
Google's compact Gemma 3 model brings the quality of a much larger system into a 4B package that runs on integrated graphics. A strong first choice for anyone wanting a capable, low-power model without needing a discrete GPU.
Min VRAM
3 GB
Context
128K
Run with Ollama
ollama run gemma3:4bGemma 3 27B
Google · Gemma Terms of Use
The largest Gemma 3 model and one of the best open-weight models in the 20–30B range. Fits in a single RTX 4090 (24 GB) with room to spare at Q4 quantization, and delivers instruction-following quality that rivals models twice its size.
Min VRAM
15 GB
Context
128K
Run with Ollama
ollama run gemma3:27bWant to mix in more filters?
Open the faceted model browser to combine VRAM, family, license, developer and task filters, then sort the results your way.