Model Collections
Curated shortlists of open-weight models you can run locally, grouped by VRAM budget, family, license and task. Pick the collection that matches your hardware or use case, then jump straight into the faceted browser. Every model runs free with Ollama.
By VRAM budget
5Best Local AI Models for 8GB VRAM
These open-weight models run comfortably on a GPU with 8 GB of VRAM at Q4 quantization, so a mid-range gaming card or a recent laptop is enough to run a real LLM locally.
Best Local AI Models for 12GB VRAM
With 12 GB of VRAM you unlock a wider set of models, including the first mixture-of-experts release, while still running everything on a single consumer card.
Best Local AI Models for 16GB VRAM
16 GB of VRAM is the sweet spot where the larger 20-30B models become practical at Q4 alongside every smaller model in the catalog.
Best Local AI Models for 24GB VRAM
A 24 GB card such as an RTX 3090 or 4090 runs nearly the whole catalog, including heavyweight coding specialists, with room to spare for longer context.
Best Local AI Models for 48GB VRAM
With 48 GB of VRAM you reach flagship 70B-class open-weight models at Q4, the kind of setup that needs a dual-GPU workstation or a 48 GB unified-memory Mac.
By model family
5Best Llama Models to Run Locally
Meta's Llama family spans the tiny 3B edge model up to the 70B flagship, covering everything from on-laptop assistants to GPT-4-class local inference.
Best Qwen Models to Run Locally
Alibaba's Qwen line includes one of the most-downloaded mid-size models of 2026 and a dedicated coding variant, all under the permissive Apache 2.0 license.
Best Gemma Models to Run Locally
Google's Gemma family runs from a 4B model on integrated graphics up to the 27B heavyweight, with some of the longest context windows in the catalog.
Best DeepSeek Models to Run Locally
DeepSeek's open-weight models bring chain-of-thought reasoning to local hardware, from a compact 8B R1 distill to the large V4 Flash mixture-of-experts, all MIT licensed.
Best Mistral Models to Run Locally
Mistral AI's models are all Apache 2.0 licensed and range from the efficient 7B workhorse to the multilingual Nemo 12B and the Mixtral 8x7B mixture-of-experts.
By license
2Commercially Free Local AI Models (Apache 2.0)
Every model here ships under the Apache 2.0 license, so you can run, modify and ship it in a commercial product with no usage restrictions.
MIT Licensed Local AI Models
These models carry the MIT license, one of the most permissive options available, which means free commercial use with virtually no strings attached.
By task
3Best Local AI Models for Coding
These models are tuned or well-suited for writing, completing and reviewing code, from lightweight autocomplete helpers to flagship agentic coders.
Best Local AI Models for Chat
Looking for a private conversational assistant? These models are well-suited for everyday chat and run entirely on your own machine.
Best Local AI Models for Reasoning
These models are built to work through multi-step math, logic and analysis problems with explicit reasoning, all running locally.
Specialized picks
3Fastest Lightweight Local AI Models
Light-tier models that ask the least of your hardware and respond fast, sorted smallest-VRAM first so you can start with the leanest option.
Mixture-of-Experts (MoE) Local AI Models
Mixture-of-experts models carry a large total parameter count but activate only a fraction per token, delivering big-model quality at a fraction of the compute.
Best Local AI Models for Apple Silicon / Mac
Apple Silicon's unified memory earns its keep on the mid-to-large models that strain a single consumer GPU. These need 16 to 48 GB at Q4 — exactly the range where a Mac's shared memory pool beats building a dual-GPU rig.