In this categoryIntegrations · 69
- How to Get a YouTube Data API Key and OAuth CredentialsStart
- How to Upload a Video to YouTube with Python and the Data API
- How to Generate YouTube Video Scripts with the Claude API
- How to Generate YouTube Titles and Descriptions with ChatGPT's API
- How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds
- How to Upload Videos to YouTube Automatically with Zapier
- How to Build a Make Scenario That Turns Topics into Draft Scripts
- How to Schedule YouTube Uploads with the publishAt Timestamp
- How to Auto-Draft Replies to YouTube Comments with Claude
- How to Bulk Update Titles and Tags Across Many Videos via API
- How to Build an End-to-End Script-to-Upload Pipeline with Claude and the API
- How to create a Telegram bot with BotFather and get your token
- How to find your Telegram chat ID for sending messages
- How to auto-post to a Telegram channel from a script
- How to make a Telegram bot reply with AI answers
- How to Build a Discord Bot That Auto-Replies With AI
- How to Build a Slack AI Bot With Bolt for Node
- How to add slash commands and a menu to a Telegram bot
- How to Auto-Reply on WhatsApp With AI Using Twilio
- How to switch a Telegram bot from polling to webhooks
- How to Add an /ask Slash Command That Calls AI in Discord
- How to schedule Telegram channel posts with cron
- How to add inline buttons and handle taps in a Telegram bot
- How to send photos, documents, and files with a Telegram bot
- How to Give Your Chat Bot Memory of the Conversation
- How to auto-post your blog RSS feed to a Telegram channel
- How to Make a Slack Command That Summarizes a Thread With AI
- How to build an AI moderator bot for a Telegram group
- How to Connect AI to WhatsApp With the Meta Cloud API
- How to Connect AI to Chat Apps Without Code Using n8n
- How to Build an AI Auto-Reply Viber Bot
- How to Stream AI Replies Live in Discord by Editing Messages
- How to Send an AI Welcome Message When Someone Opens Your Viber Bot
- How to Connect Notion to Claude Using the MCP Connector
- How to connect Gmail to ChatGPT so it drafts replies for you
- How to Add a Gemini AI Formula to Google Sheets With Apps Script
- How to use Gemini inside Gmail to summarize and reply to threads
- How to Create a Notion Internal Integration Token for AI Scripts
- How to connect Google Calendar to ChatGPT to schedule events by chat
- How to Classify Spreadsheet Rows With ChatGPT in Google Sheets
- How to Connect Google Drive to Claude and Search Your Files
- How to Use Gemini in the Google Docs Side Panel to Draft and Edit
- How to connect Outlook to AI with Power Automate to summarize emails
- How to Build a Notion to Google Sheets AI Pipeline With Zapier
- How to give Claude access to your Gmail with an MCP server
- How to Auto-Summarize Notion Meeting Notes With an AI Script
- How to auto-label and prioritize Gmail with AI using Apps Script
- How to Generate a Full Google Doc From an AI Prompt With Apps Script
- How to turn AI meeting notes into Outlook calendar follow-ups
- How to Extract and Summarize a Drive PDF With AI Using Apps Script
- How to build an AI scheduling assistant that books meetings via email
- How to Translate a Google Sheets Column With an AI Custom Function
- How to sync AI-extracted tasks from email to your calendar with n8n
- How to use AI to triage and batch-reply to your inbox each morning
- How to securely give an AI tool access to your email account
- How to Send a Zapier Webhook to OpenAI and Get a Summary Back
- How to Trigger an n8n AI Agent from a Webhook Node
- How to Classify Incoming Webhook Data with Claude in Make
- How to Auto-Draft Email Replies from a Form with Zapier and GPT
- How to Connect a Webhook to Receive Events
- How to Verify Webhook Signatures Before Calling an AI Model
- How to Give a Zapier AI Webhook Simple Conversation Memory
- How to Describe Uploaded Images with a Make Webhook and GPT Vision
- How to Add an Error Handler and Retry to an AI Webhook in Make
- How to Answer a Webhook Instantly While AI Runs in the Background in n8n
- How to Build a Slack Slash Command AI Bot with an n8n Webhook
- How to Cap AI Spend on a Webhook Across Zapier, Make, and n8n
- How to Pin an API Version for a Stable Integration
How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds
Transcribe audio with OpenAI Whisper locally on a Mac or an 8 GB NVIDIA GPU, see the actual whisper.cpp vs faster-whisper speed gap, then upload the finished SRT as YouTube captions.
OpenAI Whisper transcribes speech to text on your own machine, with no audio uploaded anywhere. The fastest local path depends on hardware: whisper.cpp with Core ML on a Mac, or faster-whisper on an NVIDIA GPU, where int8 mode fits in 8 GB of VRAM. This guide covers both, compares their speed, then uploads the transcript as YouTube captions.
What you need
- ffmpeg installed (Whisper uses it to read media)
- Python 3.9+ and the openai-whisper package for the quick-start path
- A Mac with Apple Silicon, for the whisper.cpp path, or an 8 GB+ NVIDIA GPU, for the faster-whisper path
- Your client_secret.json and the video id, if you're uploading captions to YouTube
Step 1: Install Whisper and ffmpeg
Step 2: Transcribe to SRT
The Whisper command line writes several formats. The srt output is what YouTube wants. The small model is a good balance of speed and accuracy for spoken English, though it runs on CPU only through this pip package, on any platform.
Run Whisper faster on a Mac with whisper.cpp
The pip package above runs on CPU only on a Mac, since its PyTorch backend doesn't use Apple's GPU. whisper.cpp is a from-scratch C++ port that adds Metal GPU acceleration and can dispatch to the Apple Neural Engine through Core ML, which makes it the faster local option on Apple Silicon.
| Model | M1 Pro, 8 threads | Mac Mini M1, 4 threads |
|---|---|---|
| tiny | 102 ms | 194 ms |
| base | 220 ms | 380 ms |
| small | 685 ms | 1,249 ms |
| medium | 1,928 ms | 3,980 ms |
| large | 3,350 ms | 7,979 ms |
These are encoder run times, not full transcription time, pulled from community-submitted results on the whisper.cpp benchmark issue. They're not a controlled lab test, but the gap between sizes is consistent: medium runs roughly 3 times slower than small, and large is roughly double medium again.
Run Whisper on an 8 GB GPU with faster-whisper
faster-whisper reimplements Whisper on CTranslate2, a fast inference engine, and its maintainers publish a benchmark that lands squarely in the budget-GPU range: an RTX 3070 Ti, an 8 GB card, transcribing a 13-minute test clip with the large-v2 model at beam size 5.
| Engine | Precision | Time (13-min clip) | VRAM |
|---|---|---|---|
| openai/whisper | fp16 | 2m 23s | 4,708 MB |
| whisper.cpp (Flash Attention) | fp16 | 1m 05s | — |
| faster-whisper | fp16 | 1m 03s | 4,525 MB |
| faster-whisper | int8 | 59s | 2,926 MB |
| faster-whisper, batched (batch 8) | int8 | 16s | — |
This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.
from faster_whisper import WhisperModel
model = WhisperModel("large-v2", device="cuda", compute_type="int8")
segments, info = model.transcribe("final.mp3")
for segment in segments:
print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")whisper.cpp vs faster-whisper: which is actually faster
On the numbers above, faster-whisper's int8 mode is the fastest single option overall, beating whisper.cpp's own CUDA path on the same 13-minute clip (59 seconds versus 1 minute 5 seconds) while using well under half the VRAM of the original Whisper. But faster-whisper's CTranslate2 backend has no Apple GPU support, so none of that speed carries over to a Mac. The split is hardware, not preference:
| Your setup | Pick |
|---|---|
| NVIDIA GPU, 8 GB or more | faster-whisper, int8 compute type |
| Apple Silicon Mac | whisper.cpp, with Metal or Core ML |
| CPU only, no GPU | whisper.cpp, or faster-whisper's int8 mode |
Step 3: Upload the caption track
Use the captions.insert endpoint of the YouTube Data API, passing the video id and the SRT as the media body. Reuse the service() helper from the upload guide for authentication.
from googleapiclient.http import MediaFileUpload
from upload import service # reuse the auth helper
def add_caption(video_id, srt_path, language="en"):
body = {
"snippet": {
"videoId": video_id,
"language": language,
"name": "English (Whisper)",
"isDraft": False,
}
}
media = MediaFileUpload(srt_path, mimetype="application/octet-stream")
res = service().captions().insert(part="snippet", body=body, media_body=media).execute()
print("Caption track id:", res["id"])
if __name__ == "__main__":
add_caption("dQw4w9WgXcQ", "final.srt")Result
The video now shows an English (Whisper) caption track in the CC menu. Run the same flow per language by translating the SRT and changing the language code.
Whisper handles the listening half of a local audio pipeline. For the speaking half, run Kokoro-82M or clone a voice with Breeze TTS 2, both locally, both on the same GPU you just set up for transcription.
FAQ
Is whisper.cpp faster than the original Whisper?
Yes, especially on a Mac. whisper.cpp adds Metal GPU acceleration and Core ML dispatch to the Apple Neural Engine, neither of which the original PyTorch-based openai-whisper package uses on Apple Silicon.
Do I need a GPU to run Whisper locally?
No. The pip package and whisper.cpp both run on CPU, just slower. A GPU is worth adding once you're transcribing regularly; an 8 GB NVIDIA card is enough for faster-whisper's int8 mode at any Whisper model size up to large-v2.
Can an 8 GB GPU run Whisper large-v3?
Yes, with faster-whisper's int8 quantization. In the benchmark above, large-v2 at int8 used under 3 GB of VRAM, leaving plenty of headroom on an 8 GB card; large-v3 shares the same architecture and a similar footprint.
Related guides
How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
Kokoro-82M is an 82 million parameter, fully Apache 2.0 text-to-speech model with 54 voices across 8 languages. At that size the weights are small enough to run on a CPU or almost any GPU, and installing it is one pip command.
How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
Breeze TTS 2 is a 3 billion parameter open-weight speech model that clones a voice from one reference clip and speaks English or Chinese in under 40 ms to first audio.
How to Turn a Long Video into Shorts with AI in CapCut
Feed a long recording into CapCut's Long video to shorts tool and its AI finds the strongest moments, then outputs several vertical, captioned clips automatically.
Watch related tutorials
24:16
1:42:18
28:14
41:09
9:47
8:23New guides in your inbox
Fresh step-by-step how-to guides as we publish them. One email a week, no more.