研究大模型注意力头在不同训练中的稳定性,揭示其对AI可解释性的重要影响。
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
- 分层量化注意力头在多次独立训练中的表现一致性
- 中层注意力头最不稳定但表征差异最大,深层模型更明显
- 权重衰减显著提升稳定性,适合关注模型可解释性的研究者
在机制可解释性研究中,近期工作聚焦于Transformer的‘电路’——即稀疏的单层或多层子计算单元,可能反映人类可理解的功能。然而,这些网络电路在相同深度学习架构的不同实例间很少经过严格测试以验证其稳定性。若缺乏这种检验,报告出的电路是否普遍存在于不同实验中尚不清楚,可能仅是特定估计实例的特例,从而限制了其在安全关键场景中的可信度。本文系统研究了在日益复杂的多种规模Transformer语言模型中,跨重初始化的稳定性。我们逐层量化注意力头在独立初始化训练运行中学习表示的相似性。严谨实验表明:(1)中层注意力头最不稳定,但表征差异最大;(2)更深模型表现出更强的中深度偏差;(3)深层中不稳定的注意力头比同层其他头更具功能性重要性;(4)应用权重衰减优化能显著提升注意力头在随机初始化间的稳定性;(5)残差流相对稳定。研究结果确立了电路跨实例鲁棒性作为可扩展监督的基本前提,为人工智能系统的白盒监控可能性划定了边界。
原文摘要 · Abstract (English)
In mechanistic interpretability, recent work scrutinizes transformer "circuits" - sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet, these network circuits are rarely acid-tested for their stability across different instances of the same deep learning architecture. Without this, it remains unclear whether reported circuits emerge universally across labs or turn out to be idiosyncratic to a particular estimation instance, potentially limiting confidence in safety-critical settings. Here, we systematically study stability across-refits in increasingly complex transformer language models of various sizes. We quantify, layer by layer, how similarly attention heads learn representations across independently initialized training runs. Our rigorous experiments show that (1) middle-layer heads are the least stable yet the most representationally distinct; (2) deeper models exhibit stronger mid-depth divergence; (3) unstable heads in deeper layers become more functionally important than their peers from the same layer; (4) applying weight decay optimization substantially improves attention-head stability across random model initializations; and (5) the residual stream is comparatively stable. Our findings establish the cross-instance robustness of circuits as an essential yet underappreciated prerequisite for scalable oversight, drawing contours around possible white-box monitorability of AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。