构建200+语言方言的语音识别评测基准,推动语音技术更包容。
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- 构建涵盖200多种语言、口音和方言的多语言语音评测集。
- 最佳方案在通用多语言测试中提升23%识别准确率,降低18%字符错误率。
- 聚焦口音与方言数据,显著改善跨语言泛化能力,适合关注公平性的研究者。
近年来多语言语音识别(ASR)的进步并未在所有语言和语言变体间均匀分布。为推动最先进的多语言语音模型发展,我们推出了Interspeech 2025 ML-SUPERB 2.0挑战赛。该挑战赛构建了一个包含200多个语言、口音和方言的新型测试套件,用于评估当前最优的多语言语音模型。挑战赛还基于DynaBench搭建了在线评估服务器,支持参赛者灵活设计模型架构。共收到3个团队提交的5份参赛作品,全部优于基线模型。最佳方案在通用多语言测试集上实现23%的语音识别准确率(LID)绝对提升,并将字符错误率(CER)降低18%;在带有口音和方言的数据上,其CER下降30.2%,LID准确率提升15.7%,凸显了社区挑战在促进语音技术包容性方面的重要作用。
原文摘要 · Abstract (English)
Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new test suite that consists of data from 200+ languages, accents, and dialects to evaluate SOTA multilingual speech models. The challenge also introduces an online evaluation server based on DynaBench, allowing for flexibility in model design and architecture for participants. The challenge received 5 submissions from 3 teams, all of which outperformed our baselines. The best-performing submission achieved an absolute improvement in LID accuracy of 23% and a reduction in CER of 18% when compared to the best baseline on a general multilingual test set. On accented and dialectal data, the best submission obtained 30.2% lower CER and 15.7% higher LID accuracy, showing the importance of community challenges in making speech technologies more inclusive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。