arXiv:2605.12225cs.CL2026-05被引 2

用稀疏自编码器揭示Whisper编码器中丰富的语言层次结构

On the Interpretability of Whisper Encodings Using Sparse Autoencoders

论文配图:On the Interpretability of Whisper Encodings Using Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过稀疏自编码器解析Whisper编码器内部表示
  • 发现从语音到语义的多层级、单义性特征,高层特征更易操控
  • 适合对语音模型可解释性感兴趣的学者与工程师

尽管深度变换器模型发展迅速,其内部机制仍不清晰。近期研究集中于文本型变换器,而语音识别系统(ASR)仍缺乏探索。为此,我们利用稀疏自编码器分析Whisper编码器的内部表示,发现跨越语音与语义边界的多样化单义性特征,形成从音素到语义的层次结构,并在该层次上开展因果特征调控实验,包括跨语言调控。结果表明,高层特征的调控比低层更可靠,这种不对称性可能源于低层信息的冗余编码。整体而言,本工作表明Whisper编码器所表征的语言信息层次极为丰富,远超转录任务所需。

原文摘要 · Abstract (English)

While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.

可解释性语音模型自编码器Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。