arXiv:2509.22831cs.AIcs.CL2025-09被引 6

提出五个可泛化的机制解释维度,验证了注意力头发展轨迹的跨模型一致性。

Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research

  • 构建五维泛化框架:功能、发育、位置、关系、配置对应性
  • 1-back注意力头在不同模型中发育轨迹高度一致,大模型启动更早、峰值更高
  • 适合关注大模型机制可解释性与跨模型迁移的研究者

大语言模型(LLM)的机制解释研究日益重要,但缺乏明确原则判断某模型发现能否推广至其他模型。本文提出五个可能支持泛化的对应维度:功能(满足相同功能标准)、发育(预训练中相似阶段出现)、位置(绝对或相对位置一致)、关系(与其他组件交互方式相似)、配置(对应权重空间特定区域)。为验证该框架,分析了Pythia模型(14M、70M、160M、410M)在不同随机种子下1-back注意力头(前向注意力头)在预训练中的演化。结果表明,1-back注意力头的发育轨迹在各模型间具显著一致性,而位置一致性较弱;较大模型的种子表现出更早的激活时间、更陡的上升斜率和更高的峰值。文章还回应了潜在质疑,并主张提升机制可解释性泛化能力的关键在于建立模型构成性设计属性与涌现行为之间的映射。

原文摘要 · Abstract (English)

Research on Large Language Models (LLMs) increasingly focuses on identifying mechanistic explanations for their behaviors, yet the field lacks clear principles for determining when (and how) findings from one model instance generalize to another. This paper addresses a fundamental epistemological challenge: given a mechanistic claim about a particular model, what justifies extrapolating this finding to other LLMs -- and along which dimensions might such generalizations hold? I propose five potential axes of correspondence along which mechanistic claims might generalize, including: functional (whether they satisfy the same functional criteria), developmental (whether they develop at similar points during pretraining), positional (whether they occupy similar absolute or relative positions), relational (whether they interact with other model components in similar ways), and configurational (whether they correspond to particular regions or structures in weight-space). To empirically validate this framework, I analyze "1-back attention heads" (components attending to previous tokens) across pretraining in random seeds of the Pythia models (14M, 70M, 160M, 410M). The results reveal striking consistency in the developmental trajectories of 1-back attention across models, while positional consistency is more limited. Moreover, seeds of larger models systematically show earlier onsets, steeper slopes, and higher peaks of 1-back attention. I also address possible objections to the arguments and proposals outlined here. Finally, I conclude by arguing that progress on the generalizability of mechanistic interpretability research will consist in mapping constitutive design properties of LLMs to their emergent behaviors and mechanisms.

机制解释泛化性注意力头大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。