arXiv:2602.13524cs.LGcs.AI2026-02被引 1

注意力头的奇异向量能对齐特征,为理解语言模型内部机制提供新方法。

Singular Vectors of Attention Heads Align with Features

  • 通过理论分析与实验验证,证明奇异向量在特定条件下会与特征对齐。
  • 在可直接观测特征的模型中,奇异向量对齐度高达0.92以上。
  • 提出稀疏注意力分解作为对齐的可检验指标,适用于真实模型分析。

理解语言模型中的特征表示是机械可解释性研究的核心任务。近期多项研究表明,在某些情况下可通过注意力矩阵的奇异向量推断出特征表示,但缺乏充分理论依据。本文系统回答该问题:奇异向量何时、为何与特征对齐?首先,在特征可直接观测的模型中,我们验证了奇异向量与特征存在稳健对齐(相关系数>0.92)。随后从理论上证明,这种对齐在多种假设条件下均预期成立。最后探讨实际模型中如何识别此类对齐,提出稀疏注意力分解作为可检验预测,并在真实模型中发现其行为符合预测结果。综合表明,奇异向量与特征的对齐可作为语言模型特征识别的可靠且理论支持的依据。

原文摘要 · Abstract (English)

Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the observation that feature representations can be inferred in some cases from singular vectors of attention matrices. However, sound justification for this phenomenon is lacking. In this paper we address that question, asking: why and when do singular vectors align with features? First, we demonstrate that singular vectors robustly align with features in a model where features can be directly observed. We then show theoretically that such alignment is expected under a range of conditions. We close by asking how, operationally, alignment may be recognized in real models where feature representations are not directly observable. We identify sparse attention decomposition as a testable prediction of alignment, and show evidence that it emerges in real models in a manner consistent with predictions. Together these results suggest that alignment of singular vectors with features can be a sound and theoretically justified basis for feature identification in language models.

可解释性注意力机制特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。