arXiv:2603.00941cs.CLcs.SD2026-03中稿 · ICASSP 2026被引 4

改进印度语语音识别评估,让评分更贴近真实使用体验。

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

  • 引入正字法变体捕捉框架,基于大模型生成合理拼写变化。
  • 新指标OIWER使错误率降低6.3点,模型差距缩小至11.5点。
  • 更适合低资源印度语,评估结果更接近人工听感判断。

评估印度语语音识别系统面临拼写差异、词尾拆分灵活及混合语言中非标准拼写的挑战。传统词错误率(WER)常比用户实际感知表现更差。为更贴近真实性能,需捕捉可接受的正字法变体,这对低资源语言尤为困难。本文利用大模型最新进展,提出一种生成允许拼写变化的基准框架。大量实验表明,考虑正字法变体的OIWER指标平均降低6.3个错误率点,模型性能差距显著缩小(如Gemini-Canary差距从18.1降至11.5),且较WER-SN更接近人工感知(提升4.9点)。

原文摘要 · Abstract (English)

Evaluating ASR systems for Indian languages is challenging due to spelling variations, suffix splitting flexibility, and non-standard spellings in code-mixed words. Traditional Word Error Rate (WER) often presents a bleaker picture of system performance than what human users perceive. Better aligning evaluation with real-world performance requires capturing permissible orthographic variations, which is extremely challenging for under-resourced Indian languages. Leveraging recent advances in LLMs, we propose a framework for creating benchmarks that capture permissible variations. Through extensive experiments, we demonstrate that OIWER, by accounting for orthographic variations, reduces pessimistic error rates (an average improvement of 6.3 points), narrows inflated model gaps (e.g., Gemini-Canary performance difference drops from 18.1 to 11.5 points), and aligns more closely with human perception than prior methods like WER-SN by 4.9 points.

语音识别正字法印度语评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。