Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time
Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation. The model has about 100 million parameters, and its weights are freely available . It can tell apart up to eight speakers and detect when multiple people talk at the same time. More participants, heavy background noise, or reverb push error rates higher. Paired with a speech recognition system like Parakeet , the model can produce transcripts with speaker labels, though only anonymous ones like "speaker_2." It works with both recordings and live audio.
The audio buffer can be set to four levels ranging from 30.4 down to 0.32 seconds. Shorter buffers generally reduce accuracy. On the Diarization-Bench from VoiceArena , the model currently sits in first place with a 14.72 percent error rate, ahead of the next best system at 19.3 percent. The benchmark is strict. Overlapping speech counts, and even tiny misalignments at speaker transitions are scored as errors. Compared to its predecessor, Streaming Sortformer, the new model cuts the error rate by an average of 41 percent across eight test scenarios when using a 1.04-second buffer.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.