Best Local LLM for 12GB VRAM
With 12 GB of VRAM the standout pick is OpenAI's gpt-oss-20b, a 20B mixture-of-experts model that needs about 12 GB at Q4 and generates at nearly 7B speed because only 3.6B parameters are active per token. Everything smaller, including Qwen3 8B and Gemma 4 12B, fits with room to spare.
Model data last verified on June 27, 2026. All VRAM figures are for the Q4 quantization Ollama serves by default.
Every model that fits 12 GB of VRAM
11 of the 16 models in our verified catalog fit this budget at Q4, sorted with the largest fit first. Click any model for its full spec sheet, hardware cross-reference and run commands.
| Model | Params | Quant | Min VRAM | Context | Fit |
|---|---|---|---|---|---|
| OpenAI gpt-oss-20bHistoric first | 20B total / 3.6B active (MoE) | Q4 | 12 GB | 128K | 0 GB free |
| Gemma 4 12BBest for fine-tuning | 12B | Q4 | 8 GB | 256K | 4 GB free |
| Phi-4 14BBest small coder | 14B | Q4 | 8 GB | 16K | 4 GB free |
| Mistral Nemo 12B | 12B | Q4 | 7 GB | 128K | 5 GB free |
| Llama 3.1 8B | 8B | Q4 | 5 GB | 128K | 7 GB free |
| Qwen3 8BMost popular | 8B | Q4 | 5 GB | 128K | 7 GB free |
| DeepSeek R1 Distill 8BMost private | 8B | Q4 | 5 GB | 128K | 7 GB free |
| Mistral 7B v0.3 | 7B | Q4 | 4 GB | 32K | 8 GB free |
| Qwen 2.5 Coder 7B | 7B | Q4 | 4 GB | 128K | 8 GB free |
| Gemma 3 4B | 4B | Q4 | 3 GB | 128K | 9 GB free |
| Llama 3.2 3BFastest | 3B | Q4 | 2 GB | 128K | 10 GB free |
Run the top picks with Ollama
One command each. Ollama pulls the Q4_K_M build by default and exposes an OpenAI-compatible endpoint at localhost:11434/v1.
ollama run gpt-oss:20bollama run gemma4:12bollama run phi4:14bHardware that gives you 12 GB
Tracked cards and Macs with at least 12 GB of fast memory, cheapest first.
Frequently asked questions
What is the best local LLM for 12GB of VRAM?
OpenAI's gpt-oss-20b. It needs about 12 GB at Q4, ships under Apache 2.0, and its mixture-of-experts design keeps only 3.6B parameters active per token, so it generates tokens at nearly 7B speed despite 20B total parameters. Pull it with ollama run gpt-oss:20b.
Is gpt-oss-20b a tight fit on a 12GB card?
Yes, it uses essentially the whole budget: about 12 GB at Q4, plus it wants at least 16 GB of system RAM. Close other GPU-heavy apps while it runs, or drop to Qwen3 8B (about 5 GB) if you need headroom.
Is the RTX 4070 Super good for local AI?
Yes. The RTX 4070 Super 12GB is the fastest card in its price range for 7B to 13B models, with 504 GB/s of memory bandwidth. Its 12 GB limit means 34B-class models need heavy offload, so step up to a 16 GB card if you want those.
What can I not run with 12GB of VRAM?
Anything that needs more than 12 GB at Q4: Gemma 3 27B (about 15 GB), Ornith 1.0 35B (about 20 GB) and Mixtral 8x7B (about 26 GB). For those, see the 16 GB and 24 GB tiers.
Different budget or use case?
The faceted model browser combines VRAM, family, license and task filters over the same verified catalog.