Open ASR Leaderboard Adds Its First Indian-Language Speech Recognition Test
Voice Arena and Hugging Face have added Hindi and Indian English evaluation sets to the Open ASR Leaderboard, exposing speech-recognition errors that a single overall accuracy score can hide.
Voice Arena and Hugging Face have added two new evaluation sets, called Monsoon en-IN and Monsoon hi-IN, to the Open ASR Leaderboard, which ranks automatic speech recognition (ASR) systems. Hindi, spoken by more than half a billion people, is the first added to a multilingual leaderboard tab that until now covered only European languages.
The leaderboard's usual measure, word error rate (WER), reduces a system's performance to a single number. Past research has found that ASR error rates are not spread evenly: one study found commercial systems were roughly twice as inaccurate for Black speakers as for white speakers, and another found further gaps by gender, age and accent. Those disparities do not show up on a leaderboard because standard test sets record what was said but almost nothing about who said it.
The new Monsoon sets were built to vary along nine factors β including geography, age, gender, vocabulary, devices and acoustic environment β so that errors tied to a particular group of speakers can be detected rather than averaged away. Together, the four public and private splits (two each for Hindi and Indian English) comprise 4,888 speakers, with 12 attributes recorded for each speaker.
The Indian English data was collected from 428 native districts across 24 states and union territories, and the Hindi data from up to 295 districts. In both languages, more than half of all speakers appear in only a single audio clip, and the ten largest contributors account for no more than 6.8% of total recorded duration in any split β meaning no individual voice can dominate a system's score.
Terms explained
The story so far
- Tencent Releases Open-Source Hy4 Preview AI Model With 770 Billion Parameters
- How an MIT Research Project Became the Julia Programming Language
- AI Screens 150 Million Compositions to Find New Lead-Free Electronics Materials
- Scientists Challenge the 70-Year-Old 'Lizard Brain' Myth
- Hugging Face Releases @huggingface/kernels, a Library of 207 WebGPU Kernels for Browser AI
- New Open-Source AI Model NeoMME Reads Text and Images With One Shared Network
- Open ASR Leaderboard Adds Its First Indian-Language Speech Recognition Test
