arXiv:2410.08069cs.LGcs.AI2024-10ICLR被引 1

提出新方法通过反向学习消除偏差,生成更可靠的特征重要性图。

Unlearning-based Neural Interpretations

  • 用反向学习方向动态调整输入基线,避免静态方法引入的噪声偏差。
  • 能有效擦除显著特征,使决策边界局部平滑,提升解释稳定性。
  • 适合需要可信、抗干扰模型解释的研究者使用。

基于梯度的解释方法通常需要参考点以避免特征重要性计算饱和。现有基线多采用静态函数(如常数映射、平均或模糊),会引入有害的颜色、纹理或频率假设,偏离模型实际行为,导致不规则梯度累积,使归因图存在偏差、脆弱且易被操控。我们提出UNI方法,通过将输入沿反向学习方向的最陡上升方向扰动,构建可学习、去偏且自适应的基线。该方法能发现可靠基线并成功擦除显著特征,从而在局部平滑高曲率决策边界。分析表明,反向学习是生成忠实、高效、鲁棒解释的有前景方向。

原文摘要 · Abstract (English)

Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static functions--constant mapping, averaging or blurring--inject harmful colour, texture or frequency assumptions that deviate from model behaviour. This leads to accumulation of irregular gradients, resulting in attribution maps that are biased, fragile and manipulable. Departing from the static approach, we propose UNI to compute an (un)learnable, debiased and adaptive baseline by perturbing the input towards an unlearning direction of steepest ascent. Our method discovers reliable baselines and succeeds in erasing salient features, which in turn locally smooths the high-curvature decision boundaries. Our analyses point to unlearning as a promising avenue for generating faithful, efficient and robust interpretations.

模型解释反向学习归因图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。