arXiv:2605.03671cs.CL2026-05被引 2

提出新范式,让语音识别错误率更贴近人类感知。

A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition

论文配图:A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition
图 1 · 摘自论文原文
  • 用最小编辑距离重构错误率,实现可解释性
  • 错误严重性评估与人类感知对齐,提升可读性
  • 适合关注评测指标可信度的研究者

自动语音识别中最常用的评价指标——词错误率(WER)和字符错误率(CER),因与人类感知相关性差、忽略语言和语义信息而受到广泛批评。尽管已有基于嵌入的度量试图逼近人类判断,但其结果难以解释。本文提出一种新范式,将任意选定的评价指标融入其中,生成等效的最小编辑距离(minED)。该方法使转录错误与人类感知相匹配,并首次从人类视角系统分析错误严重性,显著提升了指标的可解释性与实用性。

原文摘要 · Abstract (English)

The most commonly used metrics for evaluating automatic speech transcriptions, namely Word Error Rate (WER) and Character Error Rate (CER), have been heavily criticized for their poor correlation to human perception and their inability to take into account linguistic and semantic information. While metric-based embeddings, seeking to approximate human perception, have been proposed, their scores remain difficult to interpret, unlike WER and CER. In this article, we overcome this problem by proposing a paradigm that consists in incorporating a chosen metric into it in order to obtain an equivalent of the error rate: a Minimum Edit Distance (minED). This approach parallels transcription errors with their human perception, also allowing an original study of the severity of these errors from a human perspective.

语音识别评估指标可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。