构建印度真实语境语音识别基准,覆盖15种语言与方言差异。
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

- 基于36691人、536小时未剪辑电话对话数据,覆盖139个区域。
- 揭示方言拼写变异与音质、设备类型对识别准确率的影响。
- 适合研究真实场景下印地语系语音识别的开发者和评估者。
现有印地语语音识别(ASR)基准多采用脚本化、干净语音,并以单参考词错误率(WER)为评价标准,导致模型过度拟合特定数据集。此外,严格单一参考答案会惩罚印度语言中的自然拼写变体,包括混合英语词汇的非标准化拼写。为此,我们提出了「Voice of India」,一个封闭源代码的基准数据集,包含来自15种主要印度语言、覆盖139个地区集群的未剪辑电话对话。该数据集共收录306,230条语音片段,总计536小时,涉及36,691名说话人,且转录文本保留了真实的拼写差异。我们还进行了县级地理层面的性能分析,发现显著差异。同时,从音频质量、语速、性别、设备类型等维度展开详细分析,揭示当前语音识别系统在真实世界中的薄弱环节,为改进印度语系语音识别提供重要方向。
原文摘要 · Abstract (English)
Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian languages, including non standardized spellings of code-mixed English origin words. To address these limitations, we introduce Voice of India, a closed source benchmark built from unscripted telephonic conversations covering 15 major Indian languages across 139 regional clusters. The dataset contains 306230 utterances, totaling 536 hours of speech from 36691 speakers with transcripts accounting for spelling variations. We also analyze performance geographically at the district level, revealing disparities. Finally, we provide detailed analysis across factors such as audio quality, speaking rate, gender, and device type, highlighting where current ASR systems struggle and offering insights for improving real world Indic ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。