arXiv:2602.01442cs.LGcs.AI2026-02被引 1

梯度归因会误判Transformer中层的真正重要性,导致关键组件被低估。

Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers

  • 发现早期层的冗余特征(梯度膨胀)被高估,晚期层的关键计算(隐藏英雄)被低估。
  • 在序列排序任务中,梯度归因相关性降至0.27,个别种子甚至为负值。
  • 揭示了梯度归因对集体冗余的盲区,适合关注模型可解释性研究者阅读。

基于梯度的归因是机制可解释性的主流方法,但其能否在组件层面可靠追踪因果重要性仍缺乏验证。我们在两个算法任务和最多10个随机种子下进行因果评估,发现系统性、分层失败:梯度归因持续高估早期层的梯度膨胀(Gradient Bloats),低估晚期层的隐藏英雄(Hidden Heroes)。秩相关性从序列反转任务的ρ=0.72下降至序列排序任务的0.27,单个种子中甚至达到ρ=-0.18。这一失败源于一阶梯度归因无法检测集体冗余:联合消融导致的损伤是单个消融预测的14倍。因此,尽管梯度膨胀功能影响微弱,仍主导梯度排名;而消融隐藏英雄则使域外准确率下降36.4%±22.8%。这一早期特征提取与晚期计算的系统性倒置,凸显因果验证对电路级结论的必要性。

原文摘要 · Abstract (English)

Gradient-based attribution is the workhorse of mechanistic interpretability, yet whether it reliably tracks causal importance at the component level remains largely untested. We causally evaluate this assumption across two algorithmic tasks and up to 10 random seeds, uncovering a systematic, layer-wise failure: gradient attribution consistently overvalues early-layer \textbf{Gradient Bloats} and undervalues late-layer \textbf{Hidden Heroes}. Rank correlation collapses from $ρ= 0.72$ on sequence reversal to $0.27$ on sequence sorting, reaching $ρ= -0.18$ in individual seeds. This failure stems from first-order gradient attribution's inability to detect collective redundancy: joint Bloat ablation causes $14\times$ greater damage than individual results predict. Consequently, Bloats dominate gradient rankings despite negligible functional impact, while ablating Hidden Heroes destroys OOD accuracy ($-36.4\% \pm 22.8\%$). This systematic inversion of early-layer feature extraction and late-layer computation motivates causal validation as a prerequisite for circuit-level claims.

可解释性Transformer归因分析冗余

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。