揭示梯度解释中平滑与真实性的权衡,提出量化方法
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
- 构建谱框架统一分析解释的平滑性与真实性
- 发现代理模型平滑导致解释失真,形成可测量的'解释鸿沟'
- 适用于希望平衡解释清晰度与准确性的研究者
ReLU网络在视觉任务中广泛应用,但其突变特性常依赖单个像素进行预测,导致原始梯度解释噪声大、难解读。现有方法如GradCAM通过构建代理模型实现平滑,但牺牲了真实性。本文提出统一的谱分析框架,系统量化解释的平滑性、真实性及其权衡。基于该框架,我们量化并正则化ReLU网络对高频信息的贡献,提供一种原则性方法识别此权衡。分析揭示代理模型平滑会扭曲解释,产生可形式化定义与测量的“解释鸿沟”,并在不同设计选择、数据集和消融实验中验证了理论结果。
原文摘要 · Abstract (English)
ReLU networks, while prevalent for visual data, have sharp transitions, sometimes relying on individual pixels for predictions, making vanilla gradient-based explanations noisy and difficult to interpret. Existing methods, such as GradCAM, smooth these explanations by producing surrogate models at the cost of faithfulness. We introduce a unifying spectral framework to systematically analyze and quantify smoothness, faithfulness, and their trade-off in explanations. Using this framework, we quantify and regularize the contribution of ReLU networks to high-frequency information, providing a principled approach to identifying this trade-off. Our analysis characterizes how surrogate-based smoothing distorts explanations, leading to an ``explanation gap'' that we formally define and measure for different post-hoc methods. Finally, we validate our theoretical findings across different design choices, datasets, and ablations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。