In this categoryLocal AI · 48
More guides
Local AIIntermediate

How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters

Nemotron 3 Diarization is a 99 million parameter speaker diarization model from NVIDIA that figures out who spoke when, in real time or offline, for up to 8 speakers. It runs on CPU, needs no GPU, and ships a C++ runtime for native deployment. Here is how to install it and what the streaming latency tradeoffs look like.

7 minIntermediate

Nemotron 3 Diarization is a 99 million parameter open-weight model from NVIDIA that labels who spoke when in an audio recording, for up to 8 speakers, in either streaming or offline mode. It is small enough to run on a CPU, ships a native C++ runtime alongside the Python path, and can tag speakers on a live mic feed with as little as 320ms of latency.

Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.

Nemotron 3 Diarization entered the top 20 of Hugging Face's overall trending list. At time of writing it has 561 likes and 36,386 downloads on the nvidia/Nemotron-3-Diarization model page, plus 148 stars on the NVIDIA/NeMo-Speech.cpp GitHub repo that runs it natively. The top linked community Space is embedl/hfviewer (147 likes).

Specs at a glance

FieldValue
PublisherNVIDIA
Parameters99.2 million (per the HF API's safetensors metadata)
Architecture31-layer Transformer encoder with rotary positional embeddings (Sortformer)
PipelineVoice activity detection / speaker diarization
Max speakers8
Input16 kHz single-channel audio: .wav, .flac, .opus, .mp3
Output resolution10 ms frames (configurable in multiples of 10 ms)
Max audio lengthUnlimited, via chunked inference
LicenseOpenMDW 1.1
Release dateSeptember 1, 2026
GitHub stars148 (NVIDIA/NeMo-Speech.cpp)
Permutation order follows arrival time
Like the original Sortformer this model builds on, Nemotron 3 Diarization orders its output channels by each speaker's first appearance in the audio rather than a fixed speaker ID. The first voice you hear becomes speaker 1, the second new voice becomes speaker 2, and so on.

Hardware: this one does not need a GPU

At 99 million parameters, Nemotron 3 Diarization is light enough to run on a CPU for most use cases. The published inference speed benchmark used a Blackwell RTX PRO 5000, but that is for measuring throughput headroom, not a minimum requirement: a model this size has nowhere near the memory footprint of an LLM or a multi-billion-parameter TTS model. If you are already running an ASR pipeline on CPU, adding diarization on top costs very little extra.

On that RTX PRO 5000, at the offline 30.4 second latency setting with batch size 1, NVIDIA reports a real-time factor of 1340x in eager mode and 4385x with the model compiled. In plain terms: compiled, it can process over an hour of audio in under a second of GPU time. CPU throughput will be far lower than that, but for most diarization workloads, where the bottleneck is usually the audio pipeline rather than the model, it is still fast enough to run comfortably alongside other work.

Streaming vs. offline: picking a latency profile

The model ships one checkpoint that covers four latency profiles, controlled by five parameters measured in 80ms frames: SPKCACHE_LEN (how much speaker history the cache retains), FIFO_LEN (recent frames kept for context), CHUNK_LEN (frames processed per step), RIGHT_CONTEXT (future frames attached to each chunk), and UPDATE_PERIOD (how often the speaker cache refreshes). Lower latency means a shorter chunk and less right context, which trades a bit of accuracy for responsiveness.

ProfileInput latencyCHUNK_LENRIGHT_CONTEXTUPDATE_PERIOD
Offline (full accuracy)30.4 s34040300
Low latency1.04 s94222
Very low latency0.64 s62222
Ultra-low latency0.32 s31222

SPKCACHE_LEN and FIFO_LEN stay fixed at 264 across all four profiles; only the chunk size, right context, and cache update period change. Input buffer latency works out to (CHUNK_LEN + RIGHT_CONTEXT) times 80ms, not counting the time the model actually spends computing. For a live meeting assistant or a captioning overlay, the low-latency or very-low-latency profile is the realistic choice. For transcribing a recorded podcast or call, the offline profile gets you the best accuracy since there is no reason to rush it.

Accuracy: DER scores on two public benchmarks

Diarization error rate (DER) measures how much of the audio got assigned to the wrong speaker, missed entirely, or falsely flagged as speech; lower is better. NVIDIA's published numbers are from the offline 30.4 second latency configuration.

BenchmarkResult
DIHARD III, 1-4 speakers9.13% DER
DIHARD III, 5-9 speakers27.58% DER
DIHARD III, full set12.73% DER, 81.47% speaker-counting accuracy
CALLHOME-Part2, 2 speakers5.98% DER
CALLHOME-Part2, 3 speakers9.26% DER
CALLHOME-Part2, 4 speakers11.03% DER
CALLHOME-Part2, average9.10% DER

The gap between the 1-4 speaker and 5-9 speaker DIHARD scores is the pattern to notice: like most diarization models, accuracy holds up well for small groups and degrades as more voices overlap in the same recording. If your use case is one-on-one calls or small meetings, expect results close to the CALLHOME numbers. Large group recordings will land closer to the 27.58% figure.

Training data

NVIDIA trained the model on roughly 10,000 hours of real conversations plus 82,611 hours of simulated multi-talker mixtures, spanning English, Mandarin, Hindi, Kannada, Telugu, and Bengali across conversational, telephonic, meeting, and podcast audio. The heavy use of simulated mixtures is a common technique for diarization specifically, since real recordings with precise speaker-boundary labels for more than a couple of speakers are expensive to collect at scale.

Install it

For a quick native install, NVIDIA's NeMo-Speech.cpp runtime runs the model directly from the command line with no Python environment required.

bash - diarize with NeMo-Speech.cpp
after installing the runtime per the repo's install docs
$nemo-speech diarize meeting.wav
$

The same binary can tag an existing transcript with speaker labels in one pass:

bash - transcribe with speaker tags
$nemo-speech transcribe meeting.wav --diarize --json
$

For training, fine-tuning, or more control over the streaming parameters, use NVIDIA NeMo Speech instead. It needs Python 3.12 or later, Cython, and a recent PyTorch build.

bash - install NeMo Speech
$apt-get update && apt-get install -y libsndfile1 ffmpeg
$uv pip install Cython packaging
$uv pip install 'nemo-toolkit[asr]'
$
diarize.py
from nemo.collections.asr.models import SortformerEncLabelModel

diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")
diar_model.eval()

# low-latency streaming profile
diar_model.sortformer_modules.chunk_len = 9
diar_model.sortformer_modules.chunk_right_context = 4
diar_model.sortformer_modules.fifo_len = 264
diar_model.sortformer_modules.spkcache_update_period = 222
diar_model._check_streaming_parameters()

predicted_segments = diar_model.diarize(audio=["/path/to/your/audio.wav"], batch_size=1)

for segment in predicted_segments[0]:
    print(segment)
Numpy arrays need an explicit sample rate
If you pass a numpy array instead of a file path to diarize(), you must set sample_rate explicitly. The default is 16000, and the model expects 16 kHz mono audio regardless of input format.

FAQ

What is Nemotron 3 Diarization?

A 99 million parameter open-weight speaker diarization model from NVIDIA that determines who spoke when in an audio recording, for up to 8 speakers, in either streaming or offline mode.

Do I need a GPU to run it?

No. At under 100 million parameters, the model runs on CPU for most use cases. NVIDIA's benchmark numbers come from a Blackwell RTX PRO 5000, but that measures top-end throughput, not a hardware floor.

How many speakers can it handle?

Up to 8 speakers in a single recording. Accuracy is strongest with 1-4 speakers and degrades as more voices overlap, based on the published DIHARD III benchmark.

How low can the latency go?

As low as 320ms of input buffer latency in the ultra-low-latency streaming profile. The same checkpoint also supports a 30.4 second offline profile for maximum accuracy when speed does not matter.

What license is it under?

OpenMDW 1.1, NVIDIA's open model development weights license.

Where to go from here

  • Pairing diarization with text-to-speech for a full voice pipeline? Run Fish Audio S2 Pro locally covers a multilingual TTS model with inline emotion control.
  • Need voice cloning instead of speaker tagging? Run IndicF5 locally covers a 350M-parameter model built for 11 Indian languages.
  • Browse other speech and audio models: the local AI model directory lets you compare specs side by side.

Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)

Watch related tutorials

Free weekly email

Weekly local AI drops

New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.

Tags
#nemotron 3 diarization#nvidia nemotron diarization#speaker diarization model#streaming sortformer#run diarization locally#who spoke when audio