How-To
Step-by-step guides with video, screenshots and code. Set up your tools, connect integrations, fix errors and ship faster. Search or filter by category, then follow each guide one step at a time.
New guides in your inbox
Fresh step-by-step how-to guides as we publish them. One email a week, no more.
How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters
Nemotron 3 Diarization is a 99 million parameter speaker diarization model from NVIDIA that figures out who spoke when, in real time or offline, for up to 8 speakers. It runs on CPU, needs no GPU, and ships a C++ runtime for native deployment. Here is how to install it and what the streaming latency tradeoffs look like.
How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
Audio8 ASR Infinite is a 4.09 billion parameter streaming speech recognition model from Edge0 that transcribes Chinese and English audio 24/7 with a rolling KV cache, so memory and latency stay flat no matter how long the recording runs. Here is what it costs in VRAM and how to run it.
How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
Fish Audio S2 Pro is a 4.56 billion parameter text-to-speech model that covers 80+ languages and lets you control prosody and emotion with plain text tags embedded in the input. Here is what the model costs in VRAM and how to install it.
How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
Qwen3-TTS-12Hz-1.7B-VoiceDesign is a 1.92 billion parameter open weight text-to-speech model from Qwen that builds a voice from a written description instead of a reference clip, streams audio in as little as 97ms, and covers 10 languages. Here is what it takes to run it and how to design your first voice.
How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
YuE2-3B is a 3.63 billion parameter open source music model from m-a-p that writes an editable musical score alongside the audio, so you can hand it lyrics and a style prompt, get a full song back, then tweak the melody or harmony and regenerate. On the WildSongBench benchmark its best-of-8 score edges out Suno v5. Here is what it takes to run, what the benchmarks actually show, and how to generate your first song.
How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
tencent/AuK is a 1.5 billion parameter open source speech model that handles zero-shot voice cloning, instruction-guided TTS, pitch and speed editing, and audio source separation through one natural-language interface. It also depends on a separate 3 billion parameter Qwen2.5-Omni encoder at runtime, which changes the real hardware math. Here is what the specs actually are, how to install it, and how to run your first zero-shot clone.
How to Remove Vocals from a Suno or Udio Song (and Add Your Own)
Export just the instrumental stem from a Suno or Udio generation, record your own vocal over it in a DAW, and mix the two into a track that's genuinely yours.
How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds
Transcribe audio with OpenAI Whisper locally on a Mac or an 8 GB NVIDIA GPU, see the actual whisper.cpp vs faster-whisper speed gap, then upload the finished SRT as YouTube captions.
How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
Breeze TTS 2 is a 3 billion parameter open-weight speech model that clones a voice from one reference clip and speaks English or Chinese in under 40 ms to first audio.
How to Set Up LM Studio for Local AI
Download LM Studio, grab a model from the built-in hub, start the local server, and have a working private AI chat running on your own hardware in about ten minutes.
How to Write Original Lyrics with an AI Chatbot to Use in Suno
Use a chatbot to draft fresh, copyright-clean lyrics with proper section tags, then paste them straight into Suno or Udio.
How to Generate Your First Original Track in Udio
Create an account, write a clean prompt with your own lyrics, and produce your first extended song section in Udio.
How to Extend a Short Clip into a Full Song in Udio
Turn a 30-second Udio seed into a complete arrangement using Extend before and after, then stitch it into one track.
How to Control Song Structure with Section Tags in Suno and Udio
Use bracketed tags like [Verse], [Chorus], and [Bridge] to shape arrangement, dynamics, and instrumental breaks.
How to Write Style Prompts That Nail the Genre and Mood
Build precise style prompts for Suno and Udio using genre, instruments, tempo, and mood instead of artist names.
How to Check Licensing and Share Your AI Songs Safely
Understand ownership terms, keep your projects copyright-clean, and prepare a track for release on social or streaming.
How to Make Your First Text-to-Speech Voiceover in ElevenLabs
Turn a block of text into a natural-sounding voiceover MP3 using the ElevenLabs Text to Speech studio.
How to Dub a Video into Another Language with ElevenLabs Dubbing
Use the Dubbing studio to translate and re-voice a video into a new language while keeping the original speaker's tone.
How to Produce a Multi-Speaker Podcast Episode in ElevenLabs Studio
Build a scripted two-host podcast where each speaker has a distinct voice, then export the full mix.
How to Turn a Blog Article into a Narrated Podcast with ElevenLabs
Convert a long written article into a polished single-narrator audio episode ready to publish.
How to Generate Sound Effects from Text with ElevenLabs Sound Effects
Describe a sound in words and generate a short royalty-friendly effect clip for your video or podcast.
How to Transcribe and Summarize Audio with AI in n8n
Send an audio file to a speech-to-text model, then summarize the transcript with an LLM to turn meetings and voice notes into action items.
How to Auto-Transcribe Audio Files with Whisper in Make.com
Drop audio into a cloud folder and let Make transcribe it with Whisper, then store the text automatically.
How to Get a Gemini API Key from Google AI Studio
Create a free Gemini API key in Google AI Studio and store it safely as an environment variable.
How to Generate an Audio Overview Podcast in NotebookLM
Turn your sources into a two-host audio summary you can listen to, then steer the conversation toward what you care about.
How to Relight a Flat Product Photo Using Magnific Relight
Add studio-style lighting and mood to a dull product shot with Magnific's Relight tool and a reference image.
How to Repurpose a Podcast Episode Into Audiograms and Quote Cards
Turn a long audio episode into short captioned audiograms and shareable quote graphics for social feeds.
How to Add Auto Captions to a Video in CapCut
Use CapCut's Auto Captions feature to transcribe spoken audio into editable, styled subtitles in a couple of minutes.
How to Clean Up Background Noise in CapCut With AI
Use CapCut's noise reduction and voice enhancement to strip hiss and hum out of recorded dialogue.
How to Apply Studio Sound to Clean Up Audio in Descript
Use Descript's Studio Sound to remove background noise and make voice recordings sound like they came from a studio mic.
How to Remove Background Noise in Descript
Cut hiss, hum, and room noise from a recording using Descript's noise reduction controls and Studio Sound.
How to Record and Edit a Multi-Track Podcast in Descript
Set up separate speaker tracks, label them, and edit each voice independently for a clean multi-person podcast.
How to Auto Caption a Video with AI
Generate accurate, styled captions from your video's audio so clips stay watchable with the sound off.
How to Clean Up Noisy Audio with AI
Remove background hum, room echo, and hiss from a recording using AI noise reduction without making voices sound robotic.
How to Generate Royalty-Free Background Music with AI
Create a custom backing track with an AI music tool, trim it to your video length, and keep it clear for use.
How to Clean Up Voice Audio with AI Noise Removal
Run an AI noise-removal pass on a voice recording to strip hum and room tone without making the voice sound robotic.