source: Hugging Face Blog: **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**

level: technical

nvidia introduced nemotron 3 diarization, an open-weight 100m-parameter model for speaker diarization. it identifies who spoke when in audio, supporting up to eight speakers in live or recorded conversations. the model ranks first on voicearena's diarization-bench leaderboard with a 14.72% diarization error rate (der). it handles overlapping speech, chunked processing for flexible recording lengths, and customizable streaming latency. the model uses arrival-order speaker cache and fifo context for streaming inference.

on voicearena's initial diarization-bench, the model achieved 14.72% der across 139 english conversations totaling about 22 hours, compared with 19.3% for the next-ranked system, a 24% relative reduction. at 1.04-second input-buffer latency, it reduced der on all eight evaluation conditions, with relative reductions ranging from 9.0% on callhome-part2 to 65.2% on notsofar1 mhm. the unweighted mean relative reduction was 41.0%. throughput reached 15,113x rtfx at batch size 32 with torch.compile on an rtx pro 5000.

the model supports recommended input-buffer latencies of 30.4, 1.04, 0.64, and 0.32 seconds. shorter buffers reduce latency but may lower accuracy. it was trained on public and licensed speech data, including multispeaker conversations from david ai, with 21 languages. adding david ai data decreased compound der by 0.77 absolute points. the model outputs anonymous speaker labels, not real identities, and can be combined with asr for speaker-attributed transcripts. performance may degrade on very long recordings or noisy audio.

why it matters: accurate speaker diarization improves meeting transcripts, call analytics, and voice agent memory by attributing speech to the right person, even during overlaps.


source: Hugging Face Blog: **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**