arXiv:2604.27533cs.CL2026-04被引 11

用新指标评估语言模型重评分对语音识别错误的影响

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

  • 引入词性错误率和语义嵌入错误率,补充传统词错率
  • 发现语言模型重评分显著降低语法与语义错误
  • 适合关注语音识别质量细节的研究者

评估自动语音识别(ASR)系统是经典但困难且尚未解决的问题,常仅依赖词错误率(WER)。然而该指标存在诸多局限,无法深入分析转录错误。本文通过引入自然语言处理中常用的多种指标,研究语言模型在后处理重评分阶段对ASR结果的影响。特别提出两个与形态句法和语义相关的度量:1)词性错误率(POSER),突出语法层面的错误;2)嵌入错误率(EmbER),基于错误词语之间的语义距离对WER进行加权修正。这些指标揭示了语言模型在重评分过程中对转录结果的语言学贡献。

原文摘要 · Abstract (English)

Evaluating automatic speech recognition (ASR) systems is a classical but difficult and still open problem, which often boils down to focusing only on the word error rate (WER). However, this metric suffers from many limitations and does not allow an in-depth analysis of automatic transcription errors. In this paper, we propose to study and understand the impact of rescoring using language models in ASR systems by means of several metrics often used in other natural language processing (NLP) tasks in addition to the WER. In particular, we introduce two measures related to morpho-syntactic and semantic aspects of transcribed words: 1) the POSER (Part-of-speech Error Rate), which should highlight the grammatical aspects, and 2) the EmbER (Embedding Error Rate), a measurement that modifies the WER by providing a weighting according to the semantic distance of the wrongly transcribed words. These metrics illustrate the linguistic contributions of the language models that are applied during a posterior rescoring step on transcription hypotheses.

语音识别语言模型评估指标错误分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。