Step by step

How-To

Step-by-step guides with video, screenshots and code. Set up your tools, connect integrations, fix errors and ship faster. Search or filter by category, then follow each guide one step at a time.

550Guides
550Matching
16Categories
Free weekly email

New guides in your inbox

Fresh step-by-step how-to guides as we publish them. One email a week, no more.

Local AI
Intermediate7 min

How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters

Nemotron 3 Diarization is a 99 million parameter speaker diarization model from NVIDIA that figures out who spoke when, in real time or offline, for up to 8 speakers. It runs on CPU, needs no GPU, and ships a C++ runtime for native deployment. Here is how to install it and what the streaming latency tradeoffs look like.

Local AI
Intermediate7 min

How to Run Freya-TTS Locally: A 183M-Parameter Turkish Text-to-Speech Model

Freya-TTS (FreyaTTS-small) is a 183 million parameter open-weight Turkish text-to-speech model built on a tokenizer-free flow-matching transformer. It needs under 2GB of VRAM and runs with a real-time factor around 0.10 on a consumer GPU. Here is what it costs to run and how to install it.

Local AI
Intermediate7 min

How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops

Audio8 ASR Infinite is a 4.09 billion parameter streaming speech recognition model from Edge0 that transcribes Chinese and English audio 24/7 with a rolling KV cache, so memory and latency stay flat no matter how long the recording runs. Here is what it costs in VRAM and how to run it.

Local AI
Intermediate7 min

How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents

TeleOCR is a 1.42 billion parameter open-weight vision-language model from StarDoc-AI that parses both digital PDFs and camera-photographed documents with the same weights. Here is what it scores on, what it costs in VRAM, and how to run it.

Local AI
Advanced7 min

How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks

Xing4.0-29B-A4B is a 31.22 billion parameter mixture-of-experts model from China Telecom's Xing series that only activates 4 billion parameters per token and natively handles a 256K token context. Here is what it costs in VRAM and how to actually serve it.

Local AI
Intermediate6 min

How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages

IndicF5 is a 0.35 billion parameter open weight text-to-speech model from AI4Bharat that clones a voice from a short reference clip and speaks the result in any of 11 Indian languages. Here is what it takes to install it and generate your first clip.

Local AI
Intermediate8 min

How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control

Fish Audio S2 Pro is a 4.56 billion parameter text-to-speech model that covers 80+ languages and lets you control prosody and emotion with plain text tags embedded in the input. Here is what the model costs in VRAM and how to install it.

Local AI
Intermediate7 min

How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning

ZDTaichu5.0-9B is a 9.79 billion parameter open weight vision-language model from TaichuAI that pairs a Qwen3.5-9B backbone with an NVIDIA C-RADIOv4-H vision encoder. It reads images, multiple images, and video, and leads its size class on spatial reasoning and agent benchmarks. Here is what it takes to run it and how to send it your first image.

Local AI
Intermediate7 min

How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description

Qwen3-TTS-12Hz-1.7B-VoiceDesign is a 1.92 billion parameter open weight text-to-speech model from Qwen that builds a voice from a written description instead of a reference clip, streams audio in as little as 97ms, and covers 10 languages. Here is what it takes to run it and how to design your first voice.

Local AI
Intermediate8 min

How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks

YuE2-3B is a 3.63 billion parameter open source music model from m-a-p that writes an editable musical score alongside the audio, so you can hand it lyrics and a style prompt, get a full song back, then tweak the melody or harmony and regenerate. On the WildSongBench benchmark its best-of-8 score edges out Suno v5. Here is what it takes to run, what the benchmarks actually show, and how to generate your first song.

Local AI
Intermediate8 min

MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals

MiniCPM5-2B is OpenBMB's 2.52 billion parameter model, released under Apache 2.0, and it averages 53.9 on OpenBMB's own benchmark set, ahead of Qwen3.5-4B's 51.1 despite having roughly half the parameters. Here is how to install it through vLLM, SGLang or Transformers, what the GGUF builds actually weigh, and what the benchmark table does and doesn't prove.

Local AI
Beginner7 min

BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026

google-bert/bert-base-uncased is a 110 million parameter masked-language model from 2018 that just re-entered Hugging Face's overall trending top 20. It has 48.8 million downloads, apache-2.0 licensing, and runs comfortably on a CPU. Here is how to run it, what it's actually still useful for, and why an eight-year-old model outranks this week's releases.

Local AI
Intermediate8 min

How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing

tencent/AuK is a 1.5 billion parameter open source speech model that handles zero-shot voice cloning, instruction-guided TTS, pitch and speed editing, and audio source separation through one natural-language interface. It also depends on a separate 3 billion parameter Qwen2.5-Omni encoder at runtime, which changes the real hardware math. Here is what the specs actually are, how to install it, and how to run your first zero-shot clone.

Audio & Music
Advanced11 min

How to Remove Vocals from a Suno or Udio Song (and Add Your Own)

Export just the instrumental stem from a Suno or Udio generation, record your own vocal over it in a DAW, and mix the two into a track that's genuinely yours.

Integrations
Intermediate14 min

How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds

Transcribe audio with OpenAI Whisper locally on a Mac or an 8 GB NVIDIA GPU, see the actual whisper.cpp vs faster-whisper speed gap, then upload the finished SRT as YouTube captions.

Video Editing
Intermediate9 min

How to Turn a Long Video into Shorts with AI in CapCut

Feed a long recording into CapCut's Long video to shorts tool and its AI finds the strongest moments, then outputs several vertical, captioned clips automatically.

Local AI
Intermediate14 min

GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained

Three names keep showing up when you go to download a local model: GGUF, MLX, and the newer NVFP4. They are not interchangeable, and picking the wrong one wastes memory or leaves your hardware idle. Here is what each format is, which one your machine actually wants, and how to choose in ten seconds.

Cursor
Beginner10 min

Cursor Setup Guide: From Install to Your First AI Edit

Install Cursor, bring over your VS Code settings and extensions, pick a model, and make your first AI edit with Tab, the Agent, and a project rule.

Local AI
Intermediate8 min

Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux

Stable Diffusion 3.5 Medium is Stability AI's 2B-parameter MMDiT-X text-to-image model. Here is how it differs from the SDXL and Flux coverage already on this site, how to install it with Diffusers, and the licensing catch that kicks in above $1M in revenue.

No-Code
Beginner10 min

How to Deploy a Next.js Site to Vercel in 10 Minutes

Push your Next.js project to GitHub, import it in Vercel, and click Deploy. Vercel detects the framework automatically and hands you a live vercel.app URL in about a minute, then redeploys on every push.

Local AI
Beginner10 min

How to Set Up Local AI on 8GB of VRAM

An 8GB GPU runs the 7B to 8B model class comfortably at Q4 quantization, and Phi-4 14B just barely fits. Here is exactly what to install, which quantization format to pick, and real measured tokens-per-second so you know what to expect before you download anything.

Local AI
Beginner10 min

How to Set Up Local AI on 12GB of VRAM

12GB of VRAM is the mixture-of-experts sweet spot: OpenAI's gpt-oss-20b needs about 12GB at Q4 and, because only 3.6B of its 20B parameters activate per token, generates at nearly 7B speed. Here is how to install it, which quantization it actually ships in, and cited real-world tokens-per-second.

Local AI
Beginner11 min

Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed

16GB of VRAM unlocks Gemma 3 27B, which needs about 15GB at Q4 and just fits, plus OpenAI's gpt-oss-20b with real headroom to spare. Here is how to install both, which quantization format to use, and measured tokens-per-second on the RTX 4060 Ti 16GB and RTX 5070 Ti.

Local AI
Intermediate11 min

How to Set Up Local AI on 24GB of VRAM

A 24GB card runs Ornith 1.0 35B, the top open-weight coding model on SWE-bench Verified at 75.6%, which needs about 20GB at Q4. Gemma 3 27B and everything smaller fit with headroom left for long context. Here is the install path, quantization choice, and cited RTX 3090 vs RTX 4090 speed.

Local AI
Beginner11 min

Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed

On a 16GB M4 Mac, the sweet spot is the 7B to 8B class: Qwen3 8B and Llama 3.1 8B both need about 5GB at Q4, leaving plenty of unified memory for macOS. Here is how to install Ollama or mlx-lm, which format to pick, and llama.cpp's own measured M4 speed.

Local AI
Intermediate12 min

Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed

A 48GB M4 Pro Mac mini fits Llama 3.3 70B at Q4 quantization, about 48GB, right at the ceiling, and runs Ornith 1.0 35B or Gemma 3 27B with real headroom. Here is how to install both paths, which format to pick, and what's actually measured versus what isn't for this exact chip.

Cursor
Beginner8 min

Claude Code vs Cursor in 2026: Which to Install First

Cursor and Claude Code solve different problems. Compare Setuproll's own pricing, context and editorial-score data side by side, then install the one that matches how you actually work.

Local AI
Intermediate7 min

Qwen-Image-2512: The Text-to-Image Model That Renders Real Text

Qwen-Image-2512 is Alibaba's 20B-parameter text-to-image model, and its standout feature is generating readable English and Chinese text inside images, something SDXL and Flux still struggle with. Here is how to install it with Diffusers, what its Apache 2.0 license actually allows, and why the card does not publish a VRAM number.

Local AI
Intermediate9 min

How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design

Irodori-TTS-v4.1-Anime is a free, MIT-licensed Japanese text-to-speech model fine-tuned on anime-style speech. Clone a voice from a short reference clip, or skip the clip entirely and describe the voice you want in words.

Local AI
Intermediate9 min

How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction

Breeze TTS 2 is a 3 billion parameter open-weight speech model that clones a voice from one reference clip and speaks English or Chinese in under 40 ms to first audio.

Local AI
Beginner8 min

How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model

Kokoro-82M is an 82 million parameter, fully Apache 2.0 text-to-speech model with 54 voices across 8 languages. At that size the weights are small enough to run on a CPU or almost any GPU, and installing it is one pip command.

Automation
Intermediate15 min

Extract Invoice Data with AI in n8n

A full invoice pipeline in n8n: read the PDF, pull fields and line items with an LLM against a strict schema, check the totals, and append clean rows to Google Sheets. Handles scanned invoices too.

Dify
Intermediate12 min

How to Use Ollama Local Models in Dify

Point a self-hosted Dify at a local Ollama server, register a model like Llama 3.1 or Qwen, and run your apps fully offline with no per-token cost.

Dify
Intermediate16 min

How to Build a RAG Knowledge Base in Dify

Create a Knowledge base in Dify, upload and chunk your documents, embed them, and attach retrieval to a Chatflow so the model answers from your own content.

Dify
Beginner10 min

Dify vs n8n: Which Should You Use for AI Workflows

An honest comparison of Dify and n8n: Dify is an LLM app builder with a hosted UI and API, n8n is general automation with hundreds of integrations. When to pick each, and how to combine them.

Dify
Advanced20 min

How to Self-Host Dify in Production with Docker Compose

Clone Dify, bring it up with Docker Compose, set the production env values, put it behind a reverse proxy with HTTPS, and set up backups and upgrades.

Dify
Intermediate18 min

How to Set Up an AI Code Review Workflow in Dify

Build a Dify workflow that takes a diff or code snippet and returns a structured review covering bugs, security, and style, then call it from CI or a git hook.

Dify
Beginner15 min

Dify Tutorial: Build Your First AI App (Complete Walkthrough)

A beginner walkthrough of Dify: what it is, cloud versus self-host, creating an app, picking an app type, adding a model, writing a prompt with variables, and publishing to get an API or embed.

Claude Code
Intermediate10 min

How to Add Any MCP Server to Claude Code

Learn the claude mcp add command, the difference between stdio and remote servers, where the config lives, and how to verify a server with /mcp.

Claude Code
Intermediate11 min

Best MCP Servers for Coding in 2026

A short, verified list of the MCP servers worth adding to Claude Code for real coding work, with what each one does and the exact command to install it.

Claude Code
Intermediate9 min

Context7 Alternatives: MCP Servers for Up-to-Date Docs

What Context7 does, when you might want something else, and the real MCP servers that keep Claude Code reading current documentation instead of stale guesses.

Local AI
Beginner9 min

How to Run OpenAI's gpt-oss-20b Locally with Ollama

gpt-oss-20b is OpenAI's first open-weight model under Apache 2.0, and it runs on a 16 GB machine. Here is the exact, step-by-step setup with Ollama, from install to your first chat and API call.

Local AI
Beginner6 min

gpt-oss-20b vs gpt-oss-120b: Which Open Model Should You Run?

OpenAI shipped two open-weight models. One runs on a laptop, the other needs a serious GPU. Here is exactly how they differ and which one fits your hardware and use case.

Local AI
Intermediate12 min

Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)

A current, no-nonsense map of which local models actually fit your GPU. Pick a VRAM tier, get the right GGUF model and the exact ollama pull command, plus the quant and context math that decides whether it loads or crashes.

Local AI
Intermediate12 min

Replace GitHub Copilot With a Local Coding Model (Qwen2.5-Coder + Continue.dev)

Drop the $10–19/month subscription and run your own coding assistant inside VS Code. This is the exact setup: Ollama plus Qwen2.5-Coder for chat and tab-autocomplete through the free Continue.dev extension, with honest notes on what you give up.

Local AI
Intermediate16 min

Run DeepSeek V4 Locally: Which Size Actually Fits Your Hardware (Flash vs Pro)

DeepSeek V4 is a frontier open-weight release, but the headline size and the memory you actually need are not the same number. Here is what fits a single consumer GPU, what needs a workstation, and the realistic fallback for normal hardware.

Local AI
Beginner11 min

Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)

Try local AI before you commit to Ollama or buy a GPU. WebLLM runs a real quantized model entirely inside a browser tab using WebGPU, with nothing sent to a server. This is the zero-install path, the model size limits, and the minimal code to embed it in your own page.

Local AI
Intermediate18 min

Local OCR and Document AI: Self-Host baidu Unlimited-OCR for Private Document Parsing

Invoices, receipts, contracts and scanned PDFs are exactly the documents you do not want to ship to a per-page cloud API. baidu Unlimited-OCR is a brand-new open-weight model that parses whole PDFs in one pass, on your own hardware. Here is what it is, what it runs on, and how to go from a scanned page to structured JSON.

Showing 48 of 550