How-To
Step-by-step guides with video, screenshots and code. Set up your tools, connect integrations, fix errors and ship faster. Search or filter by category, then follow each guide one step at a time.
New guides in your inbox
Fresh step-by-step how-to guides as we publish them. One email a week, no more.
How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
ZDTaichu5.0-9B is a 9.79 billion parameter open weight vision-language model from TaichuAI that pairs a Qwen3.5-9B backbone with an NVIDIA C-RADIOv4-H vision encoder. It reads images, multiple images, and video, and leads its size class on spatial reasoning and agent benchmarks. Here is what it takes to run it and how to send it your first image.
How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
Qwen3-TTS-12Hz-1.7B-VoiceDesign is a 1.92 billion parameter open weight text-to-speech model from Qwen that builds a voice from a written description instead of a reference clip, streams audio in as little as 97ms, and covers 10 languages. Here is what it takes to run it and how to design your first voice.
MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
MiniCPM5-2B is OpenBMB's 2.52 billion parameter model, released under Apache 2.0, and it averages 53.9 on OpenBMB's own benchmark set, ahead of Qwen3.5-4B's 51.1 despite having roughly half the parameters. Here is how to install it through vLLM, SGLang or Transformers, what the GGUF builds actually weigh, and what the benchmark table does and doesn't prove.
How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
tencent/AuK is a 1.5 billion parameter open source speech model that handles zero-shot voice cloning, instruction-guided TTS, pitch and speed editing, and audio source separation through one natural-language interface. It also depends on a separate 3 billion parameter Qwen2.5-Omni encoder at runtime, which changes the real hardware math. Here is what the specs actually are, how to install it, and how to run your first zero-shot clone.
How to Set Up Local AI on 8GB of VRAM
An 8GB GPU runs the 7B to 8B model class comfortably at Q4 quantization, and Phi-4 14B just barely fits. Here is exactly what to install, which quantization format to pick, and real measured tokens-per-second so you know what to expect before you download anything.
Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
On a 16GB M4 Mac, the sweet spot is the 7B to 8B class: Qwen3 8B and Llama 3.1 8B both need about 5GB at Q4, leaving plenty of unified memory for macOS. Here is how to install Ollama or mlx-lm, which format to pick, and llama.cpp's own measured M4 speed.
Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
Qwen-Image-2512 is Alibaba's 20B-parameter text-to-image model, and its standout feature is generating readable English and Chinese text inside images, something SDXL and Flux still struggle with. Here is how to install it with Diffusers, what its Apache 2.0 license actually allows, and why the card does not publish a VRAM number.
How to Use Ollama Local Models in Dify
Point a self-hosted Dify at a local Ollama server, register a model like Llama 3.1 or Qwen, and run your apps fully offline with no per-token cost.
Replace GitHub Copilot With a Local Coding Model (Qwen2.5-Coder + Continue.dev)
Drop the $10–19/month subscription and run your own coding assistant inside VS Code. This is the exact setup: Ollama plus Qwen2.5-Coder for chat and tab-autocomplete through the free Continue.dev extension, with honest notes on what you give up.