arXiv:2606.09607cs.LGcs.AI2026-06被引 1

用闭包测试验证注意力头是否构成有效电路

Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes

  • 通过共激活聚类发现注意力头群,再用因果消融验证
  • 在两个10亿参数模型中,发现的群体经闭包测试通过
  • 仅靠共激活信号不可靠,闭包测试才是确认电路的关键

可解释性研究越来越将组件群而非单个单元视为基本对象,并通过聚类共激活统计来发现它们。本文探讨这种低成本信号是否真能识别出注意力头电路。将稀疏自编码器聚类方法应用于注意力头,但改用因果消融而非重构进行验证:聚类后执行闭包测试——消融发现的社区,对比其对每个样本的损害与随机对照组。在两个密集型10亿参数模型(Pythia 1B、OLMo 1B)和两种输入分布下,所发现的社区均通过闭包测试。在混合专家模型(OLMoE-1B-7B)中,路由条件聚类虽发现显著统计信号,却未通过闭包测试——消融反而降低损失,方向错误。进一步扩展至训练过程,注意力目标选择性和参与度比例在两者间出现解耦。结论:共激活信号仅是电路假设,闭包测试才是区分假设与真实电路的标准。

原文摘要 · Abstract (English)

Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask whether such a cheap signal actually identifies an attention-head circuit. Adapting a sparse-autoencoder clustering recipe to attention heads -- but validating by causal ablation rather than reconstruction -- we cluster heads and then run a closure test: ablate the discovered community and compare per-example damage to matched-random controls. Across two dense 1B-scale models (Pythia 1B, OLMo 1B) and two input distributions, the communities pass closure. In a Mixture-of-Experts model (OLMoE-1B-7B), route-conditional clustering recovers a statistically real signal that nonetheless does not survive closure -- ablation improves loss, the wrong direction. Extending closure across training, attention-target selectivity and participation ratio decouple from function in both directions. We conclude that a cheap signal is a circuit proposal, not a confirmed circuit; closure is what separates them.

可解释性注意力机制闭包测试神经网络结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。