arXiv:2410.07400cs.CLcs.SD2024-10NAACL被引 51

建议用字符错误率取代词错误率评估多语言语音识别

Advocating Character Error Rate for Multilingual ASR Evaluation

  • 提出用字符错误率(CER)替代传统词错误率(WER)
  • 实验证明CER与人工评估更一致,尤其在复杂语言中
  • 适合多语言语音识别研究者和评测标准制定者

语音识别系统长期以英文数据集为主,采用词错误率(WER)作为主要评价指标。然而,随着多语言场景扩展,WER在形态复杂或无明确词边界的语言中表现不佳。本文指出WER的局限性,主张将字符错误率(CER)作为多语言语音识别的核心评价指标。通过在马拉雅拉姆语、英语和阿拉伯语三种语言上的真人评估,我们发现CER比WER更接近人工判断,即使对英语也是如此。为促进后续研究,我们公开了该人类评估数据集,供未来语音识别指标基准测试使用。结果表明,在多语言语音识别评估中应优先采用或至少补充使用CER,以更好反映不同语言的差异特征。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems have traditionally been evaluated using English datasets, with the word error rate (WER) serving as the predominant metric. WER's simplicity and ease of interpretation have contributed to its widespread adoption, particularly for English. However, as ASR systems expand to multilingual contexts, WER fails in various ways, particularly with morphologically complex languages or those without clear word boundaries. Our work documents the limitations of WER as an evaluation metric and advocates for the character error rate (CER) as the primary metric in multilingual ASR evaluation. We show that CER avoids many of the challenges WER faces and exhibits greater consistency across writing systems. We support our proposition by conducting human evaluations of ASR transcriptions in three languages: Malayalam, English, and Arabic, which exhibit distinct morphological characteristics. We show that CER correlates more closely with human judgments than WER, even for English. To facilitate further research, we release our human evaluation dataset for future benchmarking of ASR metrics. Our findings suggest that CER should be prioritized, or at least supplemented, in multilingual ASR evaluations to account for the varying linguistic characteristics of different languages.

语音识别评估指标多语言字符错误率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。