揭穿Grad-ECLIP伪创新,揭示模型解释的两大根本原则
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation

- 证明梯度-中间特征法与注意力法本质等价,可简化为更高效的方法
- 指出Grad-ECLIP解释结果与原模型性能严重不符,存在根本性错误
- 提出模型解释应遵循的两条核心原则,防止误导性结论
Grad-ECLIP于ICML 2024发表,被视为基于中间特征的Transformer解释新方法。本文首先表明,此类路径并非新颖;基于已有注意力机制,我们提出了完全等价于Grad-ECLIP但计算更简单的Attention-ECLIP。通过形式推导与实验验证,证明以Grad-ECLIP为代表的中间特征路径实为注意力路径的等价变体。其次,本文揭示Grad-ECLIP方法存在缺陷:其解释结果并非源自原始模型,且与模型实际表现不一致。我们分析了问题根源,并明确提出模型解释应遵循的两条基本原则,以避免类似错误。
原文摘要 · Abstract (English)
Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper demonstrates that the intermediate features-based technical route is not a novel one. Based on the existing attention-based route, we have developed Attention-ECLIP, which is completely equivalent to Grad-ECLIP but with simpler computation. Both through formal derivation and experimental validation, we prove that the intermediate feature-based route represented by Grad-ECLIP is actually an equivalent variant of the attention-based route. Next, this paper demonstrates that the Grad-ECLIP method is flawed. The model interpretation results obtained by Grad-ECLIP are not those of the original model, and the interpretation results are misaligned with the model's performance. We analyze the causes of Grad-ECLIP's flaws and propose, or rather, explicitly emphasize two fundamental principles that model interpretation should adhere to in order to avoid similar errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。