arXiv:2502.06852cs.LGcs.AI2025-02NeurIPS被引 23

提出新方法缓解梯度法电路识别中的饱和问题,提升结果可靠性。

EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification

  • 设计渐变路径追踪机制,避开梯度饱和区域
  • 在6个数据集上使电路忠实度提升最高达17.7%
  • 适合关注模型可解释性与电路发现的研究者

理解基于Transformer的语言模型内部机制仍具挑战。基于电路发现的机制可解释性旨在通过分析神经网络内部计算子图来逆向工程模型。本文重新审视现有基于梯度的电路识别方法,发现其性能受零梯度问题或饱和效应影响,导致边属性分数对输入变化不敏感,进而产生噪声大且不可靠的属性评估。为解决饱和效应,我们提出边缘属性补丁与梯度路径(EAP-GP):该方法从输入出发,自适应沿受损与干净输入梯度差的方向行进,避开饱和区域。该策略提升了属性可靠性,增强了电路识别的忠实度。我们在GPT-2 Small、GPT-2 Medium和GPT-2 XL上对6个数据集进行了评估。实验表明,EAP-GP在电路忠实度上优于现有方法,最高提升17.7%。与人工标注的基准电路对比,其精确率和召回率达到或超过此前方法,验证了其识别准确电路的有效性。

原文摘要 · Abstract (English)

Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit existing gradient-based circuit identification methods and find that their performance is either affected by the zero-gradient problem or saturation effects, where edge attribution scores become insensitive to input changes, resulting in noisy and unreliable attribution evaluations for circuit components. To address the saturation effect, we propose Edge Attribution Patching with GradPath (EAP-GP), EAP-GP introduces an integration path, starting from the input and adaptively following the direction of the difference between the gradients of corrupted and clean inputs to avoid the saturated region. This approach enhances attribution reliability and improves the faithfulness of circuit identification. We evaluate EAP-GP on 6 datasets using GPT-2 Small, GPT-2 Medium, and GPT-2 XL. Experimental results demonstrate that EAP-GP outperforms existing methods in circuit faithfulness, achieving improvements up to 17.7%. Comparisons with manually annotated ground-truth circuits demonstrate that EAP-GP achieves precision and recall comparable to or better than previous approaches, highlighting its effectiveness in identifying accurate circuits.

可解释性电路发现梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。