arXiv:2508.15882cs.SDcs.CL2025-08AAAI被引 14

用可解释性方法揭示语音识别模型内部的语义与声学信息演化机制

Beyond Transcription: Mechanistic Interpretability in ASR

  • 采用词元透镜、线性探测等方法分析语音识别模型各层信息变化
  • 发现编码器-解码器交互导致重复幻觉,声学表示中深层编码语义偏差
  • 为提升语音模型透明度与鲁棒性提供新方向,适合关注模型可信性的研究者

可解释性方法近年来在大语言模型中备受关注,有助于理解语言表征、检测错误及分析幻觉、重复等行为。然而,这些技术在自动语音识别(ASR)领域仍鲜有探索,尽管其潜力巨大。本文系统地将词元透镜、线性探测和激活修补等成熟可解释性方法应用于ASR模型,考察声学与语义信息在模型各层中的演化过程。实验揭示了此前未知的内部动态,包括导致重复幻觉的特定编码器-解码器交互,以及深埋于声学表征中的语义偏差。这些发现表明,将可解释性技术扩展至语音识别具有重要意义,为未来提升模型透明性与鲁棒性指明了新路径。

原文摘要 · Abstract (English)

Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in automatic speech recognition (ASR), despite their potential to advance both the performance and interpretability of ASR systems. In this work, we adapt and systematically apply established interpretability methods such as logit lens, linear probing, and activation patching, to examine how acoustic and semantic information evolves across layers in ASR systems. Our experiments reveal previously unknown internal dynamics, including specific encoder-decoder interactions responsible for repetition hallucinations and semantic biases encoded deep within acoustic representations. These insights demonstrate the benefits of extending and applying interpretability techniques to speech recognition, opening promising directions for future research on improving model transparency and robustness.

可解释性语音识别模型透明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。