Best Local LLM for 24GB VRAM
A 24 GB card runs Ornith 1.0 35B, the top open-weight coding model on SWE-bench Verified at 75.6%, which needs about 20 GB at Q4. Gemma 3 27B and every lighter model in the catalog fit as well, with headroom left over for longer context.
Model data last verified on June 27, 2026. All VRAM figures are for the Q4 quantization Ollama serves by default.
Every model that fits 24 GB of VRAM
13 of the 16 models in our verified catalog fit this budget at Q4, sorted with the largest fit first. Click any model for its full spec sheet, hardware cross-reference and run commands.
| Model | Params | Quant | Min VRAM | Context | Fit |
|---|---|---|---|---|---|
| Ornith 1.0 35B#1 SWE-bench open | 35B | Q4 | 20 GB | 262K | 4 GB free |
| Gemma 3 27B | 27B | Q4 | 15 GB | 128K | 9 GB free |
| OpenAI gpt-oss-20bHistoric first | 20B total / 3.6B active (MoE) | Q4 | 12 GB | 128K | 12 GB free |
| Gemma 4 12BBest for fine-tuning | 12B | Q4 | 8 GB | 256K | 16 GB free |
| Phi-4 14BBest small coder | 14B | Q4 | 8 GB | 16K | 16 GB free |
| Mistral Nemo 12B | 12B | Q4 | 7 GB | 128K | 17 GB free |
| Llama 3.1 8B | 8B | Q4 | 5 GB | 128K | 19 GB free |
| Qwen3 8BMost popular | 8B | Q4 | 5 GB | 128K | 19 GB free |
| DeepSeek R1 Distill 8BMost private | 8B | Q4 | 5 GB | 128K | 19 GB free |
| Mistral 7B v0.3 | 7B | Q4 | 4 GB | 32K | 20 GB free |
| Qwen 2.5 Coder 7B | 7B | Q4 | 4 GB | 128K | 20 GB free |
| Gemma 3 4B | 4B | Q4 | 3 GB | 128K | 21 GB free |
| Llama 3.2 3BFastest | 3B | Q4 | 2 GB | 128K | 22 GB free |
Run the top picks with Ollama
One command each. Ollama pulls the Q4_K_M build by default and exposes an OpenAI-compatible endpoint at localhost:11434/v1.
ollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUFollama run gemma3:27bollama run gpt-oss:20bHardware that gives you 24 GB
Tracked cards and Macs with at least 24 GB of fast memory, cheapest first.
Frequently asked questions
What is the best local LLM for 24GB of VRAM?
For coding, Ornith 1.0 35B: it scores 75.6% on SWE-bench Verified, one of the highest results for any open-weight coding model, and needs about 20 GB at Q4. For general chat and analysis, Gemma 3 27B at about 15 GB leaves more room for long context.
Can a 24GB card run Llama 3.3 70B?
Not on a single card at the default quantization. Llama 3.3 70B needs about 48 GB at Q4_K_M, which means a dual-GPU rig or a Mac with 48 GB of unified memory.
Can 24GB run Mixtral 8x7B?
Not quite at the default Q4_K_M build, which needs about 26 GB. It lands just over a 24 GB card, so it belongs to the 32 GB and 48 GB class of setups.
Which GPUs have 24GB of VRAM?
The RTX 3090, the RTX 4090 (around $2,755, now out of production) and AMD's RX 7900 XTX (around $1,339). All three carry the same 24 GB ceiling, so they run the same set of models; they differ in speed, price and software ecosystem.
Different budget or use case?
The faceted model browser combines VRAM, family, license and task filters over the same verified catalog.