In this categoryIntegrations · 69
Chat Apps & Bots
Docs, Sheets & Email
IntegrationsIntermediate

How to Transcribe Audio with Whisper: Mac, 8 GB GPU, and whisper.cpp vs faster-whisper Speeds

Transcribe audio with OpenAI Whisper locally on a Mac or an 8 GB NVIDIA GPU, see the actual whisper.cpp vs faster-whisper speed gap, then upload the finished SRT as YouTube captions.

14 minIntermediate

OpenAI Whisper transcribes speech to text on your own machine, with no audio uploaded anywhere. The fastest local path depends on hardware: whisper.cpp with Core ML on a Mac, or faster-whisper on an NVIDIA GPU, where int8 mode fits in 8 GB of VRAM. This guide covers both, compares their speed, then uploads the transcript as YouTube captions.

What you need

  • ffmpeg installed (Whisper uses it to read media)
  • Python 3.9+ and the openai-whisper package for the quick-start path
  • A Mac with Apple Silicon, for the whisper.cpp path, or an 8 GB+ NVIDIA GPU, for the faster-whisper path
  • Your client_secret.json and the video id, if you're uploading captions to YouTube

Step 1: Install Whisper and ffmpeg

zsh — captions
$brew install ffmpeg
$pip install openai-whisper
$

Step 2: Transcribe to SRT

The Whisper command line writes several formats. The srt output is what YouTube wants. The small model is a good balance of speed and accuracy for spoken English, though it runs on CPU only through this pip package, on any platform.

zsh — captions
$whisper final.mp4 --model small --output_format srt
[00:00.000 --> 00:04.120] Welcome back to the channel.
Writing SRT: final.srt
$
final.srt — caption file
Explorer
final.mp4
final.srt
final.srt
11
200:00:00,000 --> 00:00:04,120
3Welcome back to the channel.
4
52
600:00:04,120 --> 00:00:08,400
7Today we are fixing a wobbly chair.
Whisper produces numbered cues with start and end times.

Run Whisper faster on a Mac with whisper.cpp

The pip package above runs on CPU only on a Mac, since its PyTorch backend doesn't use Apple's GPU. whisper.cpp is a from-scratch C++ port that adds Metal GPU acceleration and can dispatch to the Apple Neural Engine through Core ML, which makes it the faster local option on Apple Silicon.

zsh — whisper.cpp on a Mac
$git clone https://github.com/ggml-org/whisper.cpp.git
$cd whisper.cpp && sh ./models/download-ggml-model.sh small
$make
$./main -m models/ggml-small.bin -f final.wav -osrt
$
ModelM1 Pro, 8 threadsMac Mini M1, 4 threads
tiny102 ms194 ms
base220 ms380 ms
small685 ms1,249 ms
medium1,928 ms3,980 ms
large3,350 ms7,979 ms

These are encoder run times, not full transcription time, pulled from community-submitted results on the whisper.cpp benchmark issue. They're not a controlled lab test, but the gap between sizes is consistent: medium runs roughly 3 times slower than small, and large is roughly double medium again.

Run Whisper on an 8 GB GPU with faster-whisper

faster-whisper reimplements Whisper on CTranslate2, a fast inference engine, and its maintainers publish a benchmark that lands squarely in the budget-GPU range: an RTX 3070 Ti, an 8 GB card, transcribing a 13-minute test clip with the large-v2 model at beam size 5.

EnginePrecisionTime (13-min clip)VRAM
openai/whisperfp162m 23s4,708 MB
whisper.cpp (Flash Attention)fp161m 05s—
faster-whisperfp161m 03s4,525 MB
faster-whisperint859s2,926 MB
faster-whisper, batched (batch 8)int816s—

This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory.

SYSTRAN/faster-whisper README, GitHub
zsh — faster-whisper
$pip install faster-whisper
$
transcribe.py
from faster_whisper import WhisperModel

model = WhisperModel("large-v2", device="cuda", compute_type="int8")
segments, info = model.transcribe("final.mp3")

for segment in segments:
    print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")
int8 is the move on an 8 GB card
The fp16 path already fits an 8 GB GPU with room to spare, but int8 drops VRAM to under 3 GB and was still the fastest single-request mode in the benchmark above, so there's no real tradeoff at this VRAM tier.

whisper.cpp vs faster-whisper: which is actually faster

On the numbers above, faster-whisper's int8 mode is the fastest single option overall, beating whisper.cpp's own CUDA path on the same 13-minute clip (59 seconds versus 1 minute 5 seconds) while using well under half the VRAM of the original Whisper. But faster-whisper's CTranslate2 backend has no Apple GPU support, so none of that speed carries over to a Mac. The split is hardware, not preference:

Your setupPick
NVIDIA GPU, 8 GB or morefaster-whisper, int8 compute type
Apple Silicon Macwhisper.cpp, with Metal or Core ML
CPU only, no GPUwhisper.cpp, or faster-whisper's int8 mode

Step 3: Upload the caption track

Use the captions.insert endpoint of the YouTube Data API, passing the video id and the SRT as the media body. Reuse the service() helper from the upload guide for authentication.

add_captions.py
from googleapiclient.http import MediaFileUpload
from upload import service  # reuse the auth helper

def add_caption(video_id, srt_path, language="en"):
    body = {
        "snippet": {
            "videoId": video_id,
            "language": language,
            "name": "English (Whisper)",
            "isDraft": False,
        }
    }
    media = MediaFileUpload(srt_path, mimetype="application/octet-stream")
    res = service().captions().insert(part="snippet", body=body, media_body=media).execute()
    print("Caption track id:", res["id"])

if __name__ == "__main__":
    add_caption("dQw4w9WgXcQ", "final.srt")
Proofread the proper nouns
Whisper is strong but still mangles brand names and people. Skim the SRT and fix any names before uploading, since captions also feed search indexing.

Result

The video now shows an English (Whisper) caption track in the CC menu. Run the same flow per language by translating the SRT and changing the language code.

Whisper handles the listening half of a local audio pipeline. For the speaking half, run Kokoro-82M or clone a voice with Breeze TTS 2, both locally, both on the same GPU you just set up for transcription.

FAQ

Is whisper.cpp faster than the original Whisper?

Yes, especially on a Mac. whisper.cpp adds Metal GPU acceleration and Core ML dispatch to the Apple Neural Engine, neither of which the original PyTorch-based openai-whisper package uses on Apple Silicon.

Do I need a GPU to run Whisper locally?

No. The pip package and whisper.cpp both run on CPU, just slower. A GPU is worth adding once you're transcribing regularly; an 8 GB NVIDIA card is enough for faster-whisper's int8 mode at any Whisper model size up to large-v2.

Can an 8 GB GPU run Whisper large-v3?

Yes, with faster-whisper's int8 quantization. In the benchmark above, large-v2 at int8 used under 3 GB of VRAM, leaving plenty of headroom on an 8 GB card; large-v3 shares the same architecture and a similar footprint.

Related guides

Watch related tutorials

Free weekly email

New guides in your inbox

Fresh step-by-step how-to guides as we publish them. One email a week, no more.

Tags
#transcribe audio with whisper#whisper local mac#whisper 8gb gpu#whisper.cpp vs faster-whisper#srt#youtube captions