A hundred million parameters, eight speakers, a third of a second
Nvidia has released Nemotron 3 Diarization, a 99.2-million-parameter model that splits a conversation into who spoke when, with open weights on Hugging Face under the OpenMDW licence. It handles up to eight speakers, works on recorded and live audio, and detects overlapping speech when people talk over each other.
Diarization is the unglamorous half of transcription. A speech recognition model turns audio into words; a diarization model decides which words belong to which person. Get it wrong and a meeting transcript attributes a decision to the person who objected to it.
The model takes 16 kHz mono audio as Mel-spectrogram features and runs them through a 31-layer Transformer encoder with rotary positional embeddings. Two memory structures track speakers: an arrival-order speaker cache holding earlier context, and a FIFO queue for recent frames.

The numbers, and whose they are
Nvidia reports a 14.72 per cent diarization error rate on VoiceArena’s Diarization-Bench, across 139 English conversations totalling roughly 22 hours, against 19.3 per cent for the second-placed system — a relative improvement of about 24 per cent. Diarization error rate counts the share of audio time given to the wrong speaker, missed, or wrongly marked as speech, so lower is better.
Against Nvidia’s own four-speaker predecessor, Streaming Sortformer v2.1, the company reports a 41 per cent average relative reduction in error across eight evaluation sets at 1.04 seconds of buffering. On DIHARD III with the full 30.4-second buffer it puts the error at 12.73 per cent, down from 19.09.
These are the vendor’s figures on a public leaderboard, published in Nvidia’s own post. They are not an independent evaluation, and the usual caution applies until someone outside the company runs the set.
Latency is the real dial
The buffer is configurable from 0.32 to 30.4 seconds, and the model is meant to be operated at one of four points: 30.4 seconds offline, 1.04 seconds for low latency, 0.64 for very low, and 0.32 as the minimum Nvidia recommends. Accuracy rises with the buffer, because more future audio is available when the model decides who was speaking.
On throughput Nvidia reports 15,113 times real time at batch size 32 with torch.compile, against 2,619 for the earlier baseline.

What it does not do
The model labels speakers anonymously. Paired with a recogniser such as Parakeet it produces a transcript marked speaker one, speaker two and so on — it does not identify anyone by name, and nothing in the release attempts voice identification.
That is a design choice worth noting at a moment when synthetic voices are in court and voice rights are being written into guidance. A model that separates speakers without claiming to recognise them sidesteps the question entirely, and the open weights mean anyone can check that for themselves.