arXiv:2606.09899cs.LGcs.AI2026-06被引 3

发现归因修补的错误根源,并提出高效修正方法。

When Attribution Patching Lies: Diagnosis and a Second-Order Correction

论文配图:When Attribution Patching Lies: Diagnosis and a Second-Order Correction
图 1 · 摘自论文原文
  • 揭示非线性下游网络是归因误差主因,非局部曲率。
  • 提出基于海塞向量积的修正方法,仅需一次反向传播。
  • 在多种模型和扰动下提升电路识别精度,适合大规模应用。

机制可解释性的核心目标是识别驱动语言模型行为的内部组件。由于重要性估计是识别神经回路的依据,系统性误差可能导致机制误判。尽管激活修补提供黄金标准因果度量,但其计算成本过高。实践中多采用基于梯度的一阶近似归因修补,但其可靠性尚不明确。本文揭示主要误差源于下游网络的非线性,而非被修补组件的局部曲率。据此提出三项实用工具:(i) 可靠性评分以检测不可信估计,(ii) 误差界量化潜在归因偏差,(iii) 海塞向量积(HVP)修正法,仅需一次额外反向传播即可消除主导误差。在五种模型家族(124M-9B参数)及随机词与自然语境(姓名替换)扰动下的评估显示,HVP是大尺度下唯一可行的二阶修正方法,而传统基线如积分梯度在此时计算不可行。对比实验表明,多步HVP变体在显著更低计算开销下达到或超过积分梯度精度,优于先前二阶基线。该改进提升了标准基准上的电路恢复保真度,支持‘筛查-标记-修复’工作流,仅对不可靠组件投入计算资源。

原文摘要 · Abstract (English)

A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior. Because these importance estimates serve as the evidence for identifying circuits, systematic errors can lead to the misidentification of the underlying mechanisms. While activation patching provides a gold-standard causal metric, its computational cost is prohibitive at scale. Practitioners instead rely on attribution patching, a gradient-based, first-order approximation whose reliability remains poorly understood. In this work, we characterize the source of this unreliability, demonstrating that the dominant error stems from the non-linearities in the downstream network rather than local curvature at the patched component. This insight yields three practical tools: (i) a reliability score to detect untrustworthy estimates, (ii) error bounds quantifying potential attribution mis-specifications, and (iii) a Hessian-vector-product (HVP) correction that eliminates the leading-order error with only one additional backward pass. In evaluations across five model families (124M-9B parameters) and both random-token and naturalistic (name-swap) perturbations, HVP is the only second-order correction feasible at larger scale, where standard baselines like Integrated Gradients become computationally prohibitive. In comparative experiments, a multi-step HVP variant matches or exceeds the accuracy of Integrated Gradients at significantly lower compute, outperforming prior second-order baselines. These improvements lead to higher-fidelity circuit recovery on standard benchmarks and support a Screen-Flag-Fix workflow that targets computational effort only toward the components flagged as unreliable.

可解释性归因分析梯度修正二阶方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。