Step by step

How-To

Step-by-step guides with video, screenshots and code. Set up your tools, connect integrations, fix errors and ship faster. Search or filter by category, then follow each guide one step at a time.

550Guides
36Matching
16Categories
Free weekly email

New guides in your inbox

Fresh step-by-step how-to guides as we publish them. One email a week, no more.

Local AI
Intermediate7 min

How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters

Nemotron 3 Diarization is a 99 million parameter speaker diarization model from NVIDIA that figures out who spoke when, in real time or offline, for up to 8 speakers. It runs on CPU, needs no GPU, and ships a C++ runtime for native deployment. Here is how to install it and what the streaming latency tradeoffs look like.

Local AI
Intermediate7 min

How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops

Audio8 ASR Infinite is a 4.09 billion parameter streaming speech recognition model from Edge0 that transcribes Chinese and English audio 24/7 with a rolling KV cache, so memory and latency stay flat no matter how long the recording runs. Here is what it costs in VRAM and how to run it.

Local AI
Intermediate8 min

How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control

Fish Audio S2 Pro is a 4.56 billion parameter text-to-speech model that covers 80+ languages and lets you control prosody and emotion with plain text tags embedded in the input. Here is what the model costs in VRAM and how to install it.

Local AI
Intermediate7 min

How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description

Qwen3-TTS-12Hz-1.7B-VoiceDesign is a 1.92 billion parameter open weight text-to-speech model from Qwen that builds a voice from a written description instead of a reference clip, streams audio in as little as 97ms, and covers 10 languages. Here is what it takes to run it and how to design your first voice.

Local AI
Intermediate8 min

How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks

YuE2-3B is a 3.63 billion parameter open source music model from m-a-p that writes an editable musical score alongside the audio, so you can hand it lyrics and a style prompt, get a full song back, then tweak the melody or harmony and regenerate. On the WildSongBench benchmark its best-of-8 score edges out Suno v5. Here is what it takes to run, what the benchmarks actually show, and how to generate your first song.

Local AI
Intermediate8 min

How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing

tencent/AuK is a 1.5 billion parameter open source speech model that handles zero-shot voice cloning, instruction-guided TTS, pitch and speed editing, and audio source separation through one natural-language interface. It also depends on a separate 3 billion parameter Qwen2.5-Omni encoder at runtime, which changes the real hardware math. Here is what the specs actually are, how to install it, and how to run your first zero-shot clone.

Audio & Music
Advanced11 min

How to Remove Vocals from a Suno or Udio Song (and Add Your Own)

Export just the instrumental stem from a Suno or Udio generation, record your own vocal over it in a DAW, and mix the two into a track that's genuinely yours.

Integrations
Intermediate14 min

How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds

Transcribe audio with OpenAI Whisper locally on a Mac or an 8 GB NVIDIA GPU, see the actual whisper.cpp vs faster-whisper speed gap, then upload the finished SRT as YouTube captions.

Local AI
Intermediate9 min

How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction

Breeze TTS 2 is a 3 billion parameter open-weight speech model that clones a voice from one reference clip and speaks English or Chinese in under 40 ms to first audio.

Local AI
Beginner10 min

How to Set Up LM Studio for Local AI

Download LM Studio, grab a model from the built-in hub, start the local server, and have a working private AI chat running on your own hardware in about ten minutes.

Audio & Music
Beginner8 min

How to Write Original Lyrics with an AI Chatbot to Use in Suno

Use a chatbot to draft fresh, copyright-clean lyrics with proper section tags, then paste them straight into Suno or Udio.

Audio & Music
Beginner7 min

How to Generate Your First Original Track in Udio

Create an account, write a clean prompt with your own lyrics, and produce your first extended song section in Udio.

Audio & Music
Intermediate8 min

How to Extend a Short Clip into a Full Song in Udio

Turn a 30-second Udio seed into a complete arrangement using Extend before and after, then stitch it into one track.

Audio & Music
Intermediate6 min

How to Control Song Structure with Section Tags in Suno and Udio

Use bracketed tags like [Verse], [Chorus], and [Bridge] to shape arrangement, dynamics, and instrumental breaks.

Audio & Music
Beginner6 min

How to Write Style Prompts That Nail the Genre and Mood

Build precise style prompts for Suno and Udio using genre, instruments, tempo, and mood instead of artist names.

Audio & Music
Intermediate8 min

How to Check Licensing and Share Your AI Songs Safely

Understand ownership terms, keep your projects copyright-clean, and prepare a track for release on social or streaming.

Audio & Music
Beginner6 min

How to Make Your First Text-to-Speech Voiceover in ElevenLabs

Turn a block of text into a natural-sounding voiceover MP3 using the ElevenLabs Text to Speech studio.

Audio & Music
Intermediate9 min

How to Dub a Video into Another Language with ElevenLabs Dubbing

Use the Dubbing studio to translate and re-voice a video into a new language while keeping the original speaker's tone.

Audio & Music
Intermediate10 min

How to Produce a Multi-Speaker Podcast Episode in ElevenLabs Studio

Build a scripted two-host podcast where each speaker has a distinct voice, then export the full mix.

Audio & Music
Beginner8 min

How to Turn a Blog Article into a Narrated Podcast with ElevenLabs

Convert a long written article into a polished single-narrator audio episode ready to publish.

Audio & Music
Beginner5 min

How to Generate Sound Effects from Text with ElevenLabs Sound Effects

Describe a sound in words and generate a short royalty-friendly effect clip for your video or podcast.

Automation
Intermediate8 min

How to Transcribe and Summarize Audio with AI in n8n

Send an audio file to a speech-to-text model, then summarize the transcript with an LLM to turn meetings and voice notes into action items.

Automation
Intermediate7 min

How to Auto-Transcribe Audio Files with Whisper in Make.com

Drop audio into a cloud folder and let Make transcribe it with Whisper, then store the text automatically.

Gemini
Beginner6 min

How to Get a Gemini API Key from Google AI Studio

Create a free Gemini API key in Google AI Studio and store it safely as an environment variable.

Gemini
Beginner6 min

How to Generate an Audio Overview Podcast in NotebookLM

Turn your sources into a two-host audio summary you can listen to, then steer the conversation toward what you care about.

Image Tools
Intermediate7 min

How to Relight a Flat Product Photo Using Magnific Relight

Add studio-style lighting and mood to a dull product shot with Magnific's Relight tool and a reference image.

Social Publishing
Intermediate8 min

How to Repurpose a Podcast Episode Into Audiograms and Quote Cards

Turn a long audio episode into short captioned audiograms and shareable quote graphics for social feeds.

Video Editing
Beginner6 min

How to Add Auto Captions to a Video in CapCut

Use CapCut's Auto Captions feature to transcribe spoken audio into editable, styled subtitles in a couple of minutes.

Video Editing
Beginner5 min

How to Clean Up Background Noise in CapCut With AI

Use CapCut's noise reduction and voice enhancement to strip hiss and hum out of recorded dialogue.

Video Editing
Beginner6 min

How to Apply Studio Sound to Clean Up Audio in Descript

Use Descript's Studio Sound to remove background noise and make voice recordings sound like they came from a studio mic.

Video Editing
Beginner6 min

How to Remove Background Noise in Descript

Cut hiss, hum, and room noise from a recording using Descript's noise reduction controls and Studio Sound.

Video Editing
Intermediate10 min

How to Record and Edit a Multi-Track Podcast in Descript

Set up separate speaker tracks, label them, and edit each voice independently for a clean multi-person podcast.

Video Editing
Beginner8 min

How to Auto Caption a Video with AI

Generate accurate, styled captions from your video's audio so clips stay watchable with the sound off.

Audio & Music
Beginner7 min

How to Clean Up Noisy Audio with AI

Remove background hum, room echo, and hiss from a recording using AI noise reduction without making voices sound robotic.

Audio & Music
Beginner8 min

How to Generate Royalty-Free Background Music with AI

Create a custom backing track with an AI music tool, trim it to your video length, and keep it clear for use.

Audio & Music
Beginner8 min

How to Clean Up Voice Audio with AI Noise Removal

Run an AI noise-removal pass on a voice recording to strip hum and room tone without making the voice sound robotic.

Showing 36 of 36