arXiv:2412.07937cs.CL2024-12被引 4

用多参考转录文本评估语音识别,发现现有错误率被高估

Style-agnostic evaluation of ASR using multiple reference transcripts

  • 采用相反风格的多参考转录文本进行无风格依赖评估
  • 发现当前SOTA语音识别系统的错误率被显著高估
  • 适合对比不同风格训练模型的质量差异

词错误率(WER)作为语音识别的评估指标存在诸多局限。评估数据集本身存在风格、正式程度和转录任务内在模糊性等问题。本文通过使用在相反风格参数下生成的多个参考转录文本,实现无风格依赖的自动语音识别(ASR)系统评估。结果表明,现有基于单参考的WER报告可能显著高估了先进ASR系统的真实内容性错误数量。此外,多参考方法可有效比较训练数据风格或目标任务不同的ASR模型质量。

原文摘要 · Abstract (English)

Word error rate (WER) as a metric has a variety of limitations that have plagued the field of speech recognition. Evaluation datasets suffer from varying style, formality, and inherent ambiguity of the transcription task. In this work, we attempt to mitigate some of these differences by performing style-agnostic evaluation of ASR systems using multiple references transcribed under opposing style parameters. As a result, we find that existing WER reports are likely significantly over-estimating the number of contentful errors made by state-of-the-art ASR systems. In addition, we have found our multireference method to be a useful mechanism for comparing the quality of ASR models that differ in the stylistic makeup of their training data and target task.

语音识别评估方法错误率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。