arXiv:2507.22047cs.AI2025-07被引 12

提升残障人士语音识别准确率,挑战赛用400小时数据设新基准

The Interspeech 2025 Speech Accessibility Project Challenge

  • 基于400小时真实残障语音数据训练与评估模型
  • 最高识别准确率达8.11% WER,语义得分88.44%
  • 适合关注无障碍语音技术的研究者与开发者

过去十年自动语音识别(ASR)系统取得显著进展,但针对言语障碍人群的性能仍不理想,部分原因在于公开训练数据有限。为此,2025年国际语音通信大会(Interspeech)启动语音可及性项目(SAP)挑战赛,使用超过400小时、来自500多位不同言语障碍者的语音数据进行收集与转录。比赛在EvalAI平台通过远程评测流程进行,以词错误率(WER)和语义得分(SemScore)为评估指标。共有22支队伍提交有效结果,其中12支在WER上优于whisper-large-v2基线,17支在语义得分上超越基线。表现最佳团队同时达到8.11%的最低WER与88.44%的最高语义得分,为未来残障语音识别系统设立了新标杆。

原文摘要 · Abstract (English)

While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training data. To bridge this gap, the 2025 Interspeech Speech Accessibility Project (SAP) Challenge was launched, utilizing over 400 hours of SAP data collected and transcribed from more than 500 individuals with diverse speech disabilities. Hosted on EvalAI and leveraging the remote evaluation pipeline, the SAP Challenge evaluates submissions based on Word Error Rate and Semantic Score. Consequently, 12 out of 22 valid teams outperformed the whisper-large-v2 baseline in terms of WER, while 17 teams surpassed the baseline on SemScore. Notably, the top team achieved the lowest WER of 8.11\%, and the highest SemScore of 88.44\% at the same time, setting new benchmarks for future ASR systems in recognizing impaired speech.

语音识别残障适配ASR挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。