现有语音识别评估方法对印地语系文字存在误导性优化,影响公平比较。
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations
- 分析主流模型文本归一化流程,发现其对印度语言脚本有系统性偏差。
- 实验证明归一化导致印地语系性能指标被人为夸大,掩盖真实缺陷。
- 建议结合本土语言学知识重构归一化流程,提升多语言评测可信度。
本文探究多语言自动语音识别(ASR)模型评估中的陷阱,聚焦印地语系文字。研究分析了OpenAI Whisper、Meta MMS、Seamless及Assembly AI Conformer等主流模型的文本归一化流程,发现其虽旨在通过去除拼写、标点和特殊字符差异以实现公平比较,但对印地语系文字存在根本性缺陷。通过文本相似度评分与深度语言学分析,证实该做法导致印地语系性能指标被人为虚高。研究呼吁发展基于母语语言学知识的归一化方案,以实现更稳健、准确的多语言ASR模型评估。
原文摘要 · Abstract (English)
This paper explores the pitfalls in evaluating multilingual automatic speech recognition (ASR) models, with a particular focus on Indic language scripts. We investigate the text normalization routine employed by leading ASR models, including OpenAI Whisper, Meta's MMS, Seamless, and Assembly AI's Conformer, and their unintended consequences on performance metrics. Our research reveals that current text normalization practices, while aiming to standardize ASR outputs for fair comparison, by removing inconsistencies such as variations in spelling, punctuation, and special characters, are fundamentally flawed when applied to Indic scripts. Through empirical analysis using text similarity scores and in-depth linguistic examination, we demonstrate that these flaws lead to artificially improved performance metrics for Indic languages. We conclude by proposing a shift towards developing text normalization routines that leverage native linguistic expertise, ensuring more robust and accurate evaluations of multilingual ASR models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。