改进跨语言语音识别评估,让不同语言的准确率更可比。
OpenWER: Improving Cross-Lingual ASR Evaluation and Enabling Token-Based Accuracy Metrics

- 用语言特异性归一化和复合词检测提升WER鲁棒性
- 52种语言测试中,最高降低25%的错误率
- 支持细粒度评分,适合多语言模型研究者
深度学习和端到端语音识别的发展催生了强大的多语言模型,但评估指标仍难以准确衡量性能。现有改进或替代常用指标词错误率(WER)的研究多集中于英语,低资源语言的评估长期被忽视,阻碍了公平的跨语言比较。我们提出OpenWER,一个开源实现,通过语言特异性归一化和复合词检测提升WER的鲁棒性,并采用基于标记的莱文斯坦对齐,保留互补指标并支持元数据嵌入以获得更细粒度的准确率评分。对52种语言的分析显示,相比常见工具库,其绝对WER降低高达25%。OpenWER通过提高多种语言下WER的可靠性,推动了语音识别研究的公平性,并支持更全面的准确性评估。
原文摘要 · Abstract (English)
Advances in deep learning and end-to-end Automatic Speech Recognition (ASR) have enabled robust multilingual models, but evaluation metrics remain limited in assessing accuracy. Efforts to improve or replace the common metric Word Error Rate (WER) often focus on English, leaving evaluations for low-resource languages under-explored and hindering fair cross-lingual comparisons. We present OpenWER, an open-source implementation that improves WER robustness through language-specific normalisation and compound word detection. A token-based Levenshtein alignment preserves complementary metrics and allows metadata embedding for granular accuracy scores. Our analysis of 52 languages shows absolute WER reductions of up to 25% compared to common libraries. OpenWER contributes to fairness in ASR research by increasing the reliability of WER across diverse languages and enabling more comprehensive accuracy evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。