source: Hugging Face Blog: The Open ASR Leaderboard Adds Its First Global South Language
level: technical
the open asr leaderboard now includes monsoon, two new evaluation sets for hindi and indian english. built with voice arena, these sets are the first global south languages on the leaderboard. they aim to expose performance gaps hidden by average word error rates. each set has public and private splits, with 4,888 speakers total and 12 speaker attributes recorded per clip. the design varies geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and multiple valid transcripts.
the indian english public set has 5.62 hours from 1,444 speakers across 428 districts and 556 device models. the hindi public set has 1.33 hours from 468 speakers across 202 districts and 315 devices. no single device model exceeds 2.1% of segments. more than half of speakers appear exactly once, so no voice dominates the score. hindi references use a lattice of accepted spellings because normalizers cannot handle the variation. eight leaderboard models score between 4.81 and 4.99 wer on indian english, but regional differences reach 0.46 wer.
the metadata includes occupation, education, income, handset brand, and years in district. this allows disaggregated analysis by region, age, or device. prior work found district-level error rates from 4% to 44% in indian asr. monsoon makes such analysis possible on a public leaderboard. the collection used peer-to-peer conversations on contributors' own phones, preserving real-world noise. quality checks included language identification, playback detection, and human transcription with five-level verification. the sets are small in hours but large in speaker diversity.
why it matters: monsoon lets developers measure speech recognition fairness across hindi and indian english speakers, revealing performance gaps that average scores hide.
source: Hugging Face Blog: The Open ASR Leaderboard Adds Its First Global South Language