In this categoryLocal AI · 48
- How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters
- How to Run Freya-TTS Locally: A 183M-Parameter Turkish Text-to-Speech Model
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters
Nemotron 3 Diarization is a 99 million parameter speaker diarization model from NVIDIA that figures out who spoke when, in real time or offline, for up to 8 speakers. It runs on CPU, needs no GPU, and ships a C++ runtime for native deployment. Here is how to install it and what the streaming latency tradeoffs look like.
Nemotron 3 Diarization is a 99 million parameter open-weight model from NVIDIA that labels who spoke when in an audio recording, for up to 8 speakers, in either streaming or offline mode. It is small enough to run on a CPU, ships a native C++ runtime alongside the Python path, and can tag speakers on a live mic feed with as little as 320ms of latency.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.
A new top-20 entrant on Hugging Face's trending list
Nemotron 3 Diarization entered the top 20 of Hugging Face's overall trending list. At time of writing it has 561 likes and 36,386 downloads on the nvidia/Nemotron-3-Diarization model page, plus 148 stars on the NVIDIA/NeMo-Speech.cpp GitHub repo that runs it natively. The top linked community Space is embedl/hfviewer (147 likes).
Specs at a glance
| Field | Value |
|---|---|
| Publisher | NVIDIA |
| Parameters | 99.2 million (per the HF API's safetensors metadata) |
| Architecture | 31-layer Transformer encoder with rotary positional embeddings (Sortformer) |
| Pipeline | Voice activity detection / speaker diarization |
| Max speakers | 8 |
| Input | 16 kHz single-channel audio: .wav, .flac, .opus, .mp3 |
| Output resolution | 10 ms frames (configurable in multiples of 10 ms) |
| Max audio length | Unlimited, via chunked inference |
| License | OpenMDW 1.1 |
| Release date | September 1, 2026 |
| GitHub stars | 148 (NVIDIA/NeMo-Speech.cpp) |
Hardware: this one does not need a GPU
At 99 million parameters, Nemotron 3 Diarization is light enough to run on a CPU for most use cases. The published inference speed benchmark used a Blackwell RTX PRO 5000, but that is for measuring throughput headroom, not a minimum requirement: a model this size has nowhere near the memory footprint of an LLM or a multi-billion-parameter TTS model. If you are already running an ASR pipeline on CPU, adding diarization on top costs very little extra.
On that RTX PRO 5000, at the offline 30.4 second latency setting with batch size 1, NVIDIA reports a real-time factor of 1340x in eager mode and 4385x with the model compiled. In plain terms: compiled, it can process over an hour of audio in under a second of GPU time. CPU throughput will be far lower than that, but for most diarization workloads, where the bottleneck is usually the audio pipeline rather than the model, it is still fast enough to run comfortably alongside other work.
Streaming vs. offline: picking a latency profile
The model ships one checkpoint that covers four latency profiles, controlled by five parameters measured in 80ms frames: SPKCACHE_LEN (how much speaker history the cache retains), FIFO_LEN (recent frames kept for context), CHUNK_LEN (frames processed per step), RIGHT_CONTEXT (future frames attached to each chunk), and UPDATE_PERIOD (how often the speaker cache refreshes). Lower latency means a shorter chunk and less right context, which trades a bit of accuracy for responsiveness.
| Profile | Input latency | CHUNK_LEN | RIGHT_CONTEXT | UPDATE_PERIOD |
|---|---|---|---|---|
| Offline (full accuracy) | 30.4 s | 340 | 40 | 300 |
| Low latency | 1.04 s | 9 | 4 | 222 |
| Very low latency | 0.64 s | 6 | 2 | 222 |
| Ultra-low latency | 0.32 s | 3 | 1 | 222 |
SPKCACHE_LEN and FIFO_LEN stay fixed at 264 across all four profiles; only the chunk size, right context, and cache update period change. Input buffer latency works out to (CHUNK_LEN + RIGHT_CONTEXT) times 80ms, not counting the time the model actually spends computing. For a live meeting assistant or a captioning overlay, the low-latency or very-low-latency profile is the realistic choice. For transcribing a recorded podcast or call, the offline profile gets you the best accuracy since there is no reason to rush it.
Accuracy: DER scores on two public benchmarks
Diarization error rate (DER) measures how much of the audio got assigned to the wrong speaker, missed entirely, or falsely flagged as speech; lower is better. NVIDIA's published numbers are from the offline 30.4 second latency configuration.
| Benchmark | Result |
|---|---|
| DIHARD III, 1-4 speakers | 9.13% DER |
| DIHARD III, 5-9 speakers | 27.58% DER |
| DIHARD III, full set | 12.73% DER, 81.47% speaker-counting accuracy |
| CALLHOME-Part2, 2 speakers | 5.98% DER |
| CALLHOME-Part2, 3 speakers | 9.26% DER |
| CALLHOME-Part2, 4 speakers | 11.03% DER |
| CALLHOME-Part2, average | 9.10% DER |
The gap between the 1-4 speaker and 5-9 speaker DIHARD scores is the pattern to notice: like most diarization models, accuracy holds up well for small groups and degrades as more voices overlap in the same recording. If your use case is one-on-one calls or small meetings, expect results close to the CALLHOME numbers. Large group recordings will land closer to the 27.58% figure.
Training data
NVIDIA trained the model on roughly 10,000 hours of real conversations plus 82,611 hours of simulated multi-talker mixtures, spanning English, Mandarin, Hindi, Kannada, Telugu, and Bengali across conversational, telephonic, meeting, and podcast audio. The heavy use of simulated mixtures is a common technique for diarization specifically, since real recordings with precise speaker-boundary labels for more than a couple of speakers are expensive to collect at scale.
Install it
For a quick native install, NVIDIA's NeMo-Speech.cpp runtime runs the model directly from the command line with no Python environment required.
The same binary can tag an existing transcript with speaker labels in one pass:
For training, fine-tuning, or more control over the streaming parameters, use NVIDIA NeMo Speech instead. It needs Python 3.12 or later, Cython, and a recent PyTorch build.
from nemo.collections.asr.models import SortformerEncLabelModel
diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")
diar_model.eval()
# low-latency streaming profile
diar_model.sortformer_modules.chunk_len = 9
diar_model.sortformer_modules.chunk_right_context = 4
diar_model.sortformer_modules.fifo_len = 264
diar_model.sortformer_modules.spkcache_update_period = 222
diar_model._check_streaming_parameters()
predicted_segments = diar_model.diarize(audio=["/path/to/your/audio.wav"], batch_size=1)
for segment in predicted_segments[0]:
print(segment)FAQ
What is Nemotron 3 Diarization?
A 99 million parameter open-weight speaker diarization model from NVIDIA that determines who spoke when in an audio recording, for up to 8 speakers, in either streaming or offline mode.
Do I need a GPU to run it?
No. At under 100 million parameters, the model runs on CPU for most use cases. NVIDIA's benchmark numbers come from a Blackwell RTX PRO 5000, but that measures top-end throughput, not a hardware floor.
How many speakers can it handle?
Up to 8 speakers in a single recording. Accuracy is strongest with 1-4 speakers and degrades as more voices overlap, based on the published DIHARD III benchmark.
How low can the latency go?
As low as 320ms of input buffer latency in the ultra-low-latency streaming profile. The same checkpoint also supports a 30.4 second offline profile for maximum accuracy when speed does not matter.
What license is it under?
OpenMDW 1.1, NVIDIA's open model development weights license.
Where to go from here
- Pairing diarization with text-to-speech for a full voice pipeline? Run Fish Audio S2 Pro locally covers a multilingual TTS model with inline emotion control.
- Need voice cloning instead of speaker tagging? Run IndicF5 locally covers a 350M-parameter model built for 11 Indian languages.
- Browse other speech and audio models: the local AI model directory lets you compare specs side by side.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
12:20
14:15
9:50Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.