arXiv:2511.20273cs.LGcs.AI2025-11NeurIPS被引 9

通过奇异向量揭示Transformer内部隐藏的子功能结构。

Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits

  • 将注意力头和MLP分解为正交奇异方向,发现内部存在多重独立计算。
  • 在IOI、GP等任务中,经典功能头包含多个重叠的子功能,对应不同方向。
  • 适合研究模型机制、可解释性与内部结构的学者参考。

基于Transformer的语言模型表现出复杂且分布式的行为,但其内部计算过程仍不清晰。现有机械可解释性方法通常将注意力头和多层感知机(MLPs)视为不可分割的整体,忽略了其中可能存在的功能性子结构。本文提出更细粒度的视角,将这些组件分解为正交的奇异方向,揭示了单个头或MLP内存在的叠加与独立计算。我们在广泛使用的标准任务如间接宾语识别(IOI)、性别代词(GP)和大于判断(GT)上验证了该方法,发现先前被识别为典型功能头(如名称移动头)编码了多个与不同奇异方向对齐的重叠子功能。计算图中此前被识别为电路元素的节点在特定低秩方向上激活强烈,表明有意义的计算存在于紧凑子空间中。尽管某些方向仍难以完全解释,结果表明Transformer的计算比以往认为的更具分布性、结构性和组合性。这一视角为细粒度机械可解释性开辟了新路径,有助于深入理解模型内部机制。

原文摘要 · Abstract (English)

Transformer-based language models exhibit complex and distributed behavior, yet their internal computations remain poorly understood. Existing mechanistic interpretability methods typically treat attention heads and multilayer perceptron layers (MLPs) (the building blocks of a transformer architecture) as indivisible units, overlooking possibilities of functional substructure learned within them. In this work, we introduce a more fine-grained perspective that decomposes these components into orthogonal singular directions, revealing superposed and independent computations within a single head or MLP. We validate our perspective on widely used standard tasks like Indirect Object Identification (IOI), Gender Pronoun (GP), and Greater Than (GT), showing that previously identified canonical functional heads, such as the name mover, encode multiple overlapping subfunctions aligned with distinct singular directions. Nodes in a computational graph, that are previously identified as circuit elements show strong activation along specific low-rank directions, suggesting that meaningful computations reside in compact subspaces. While some directions remain challenging to interpret fully, our results highlight that transformer computations are more distributed, structured, and compositional than previously assumed. This perspective opens new avenues for fine-grained mechanistic interpretability and a deeper understanding of model internals.

可解释性Transformer奇异向量机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。