arXiv:2607.05612cs.CLeess.AS2026-07

发现现代语音识别中语言模型困惑度与错误率关系不再线性,需考虑模型内部语言建模能力。

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

  • 通过剥离模型内部语言建模能力,重新检验困惑度与错误率关系。
  • 现代端到端系统中,困惑度与错误率的线性关系已被打破。
  • 使用大语言模型时需注意内部语言建模对评估结果的影响,适合语音识别研究者阅读。

语言模型困惑度(PPL)长期以来被用作自动语音识别(ASR)词错误率(WER)的代理指标,以往研究指出两者在对数空间中呈近似线性关系。然而,现代端到端ASR系统挑战了这一假设:它们已具备内部语言建模能力,通常不依赖外部语言模型,并可通过不同策略结合神经语言模型和大型语言模型(LLMs)。本文重新考察现代ASR系统中PPL与WER的关系,研究外部语言模型是否仍能提升性能、该关系在对数空间中是否保持线性、编码器上下文长度的影响,以及大语言模型困惑度是否符合传统趋势。此外,本文分析基于注意力机制的编码器-解码器系统中的内部语言建模(ILM),并证明扣除ILM后,观测到的PPL-WER关系发生变化,表明在评估外部语言模型质量时必须考虑解码器的内部语言建模能力。

原文摘要 · Abstract (English)

Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately linear relation in log-log space. Modern end-to-end ASR systems challenge this assumption because they already contain internal language modeling capacity, are often evaluated without external language models, and can now be combined with neural LMs and large language models (LLMs) through different recognition strategies. This paper revisits the relation between PPL and WER for modern ASR systems. We study whether external LMs still improve current end-to-end ASR systems, whether the PPL-WER relation remains linear in log-log space, how encoder context length affects this relation, and how LLM perplexities fit into the trend observed for standard neural LMs. We further investigate internal language modeling (ILM) in attention-based encoder-decoder systems and show that ILM subtraction changes the observed PPL-WER relation, indicating that the decoder's internal LM must be considered when interpreting the effect of external LM quality.

语音识别语言模型困惑度端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。