arXiv:2602.17598cs.CLcs.AI2026-02被引 3

揭示语音大模型本质是伪端到端,实为隐式转录流水线。

The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?

  • 通过匹配骨干测试分离语音模型与推理能力,验证其行为机制。
  • 发现语音模型内部显式出现文字转录,且文本表征具因果必要性。
  • 在噪声环境下性能反超消失,实际成本高于传统级联方案。

语音大模型普遍被认为优于传统ASR→LLM级联结构,因其直接处理音频而非仅依赖转录文本。本文提出一种评估方法与机制解释框架。首先引入匹配骨干测试,分离语音大模型的语音理解能力与底层语言模型的推理能力。其次,利用logit lens和LEACE进行机制分析,发现模型隐藏层中存在显式的文字转录信号,且文本表示具有因果必要性。此外,研究显示当前多数部署场景下,语音大模型实为高成本级联系统,在噪声环境下表现更差,清洁条件下优势在0dB信噪比时最多下降7.6%。

原文摘要 · Abstract (English)

Speech LLMs are widely understood to be better than ASR$\rightarrow$LLM cascades since they have access to the audio directly, and not just the transcript. In this paper, we present an evaluation methodology and a mechanistic interpretation of the observed behavior of speech LLMs. First, we introduce matched-backbone testing which separates out the behavior of the speech LLM from the reasoning capabilities of the underlying LLM. Second, we provide a mechanistic analysis of speech LLMs using logit lens and LEACE and show the literal transcript emerging from the LLM's hidden states and that text representations are causally necessary. We also show that in most deployed use cases, current speech LLMs are expensive cascades, and under noise, they are worse ones, with clean-condition advantages reversing by up to 7.6% at 0dB.

语音大模型机制分析端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。