arXiv:2512.06665cs.LGcs.AI2025-12

提出新评估方法,更准确衡量特征归因的鲁棒性

Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods

  • 基于生成对抗网络构造相似输入,改进鲁棒性评估
  • 新指标能揭示归因方法缺陷而非模型本身问题
  • 适合研究可解释性与模型可信度的学者参考

本文研究深度神经网络中特征归因方法的鲁棒性。针对现有归因鲁棒性评价忽视模型输出差异的问题,提出新的相似输入定义、鲁棒性度量方法,并设计基于生成对抗网络的输入生成技术。通过对比现有指标与前沿归因方法的全面评估,发现亟需一种更客观的指标,以暴露归因方法本身的弱点而非神经网络的缺陷,从而实现对归因方法鲁棒性的更准确评估。

原文摘要 · Abstract (English)

This paper studies the robustness of feature attribution methods for deep neural networks. It challenges the current notion of attributional robustness that largely ignores the difference in the model's outputs and introduces a new way of evaluating the robustness of attribution methods. Specifically, we propose a new definition of similar inputs, a new robustness metric, and a novel method based on generative adversarial networks to generate these inputs. In addition, we present a comprehensive evaluation with existing metrics and state-of-the-art attribution methods. Our findings highlight the need for a more objective metric that reveals the weaknesses of an attribution method rather than that of the neural network, thus providing a more accurate evaluation of the robustness of attribution methods.

可解释性特征归因鲁棒性GAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。