提出对抗性掩码方法,更准确评估图像音频等任务中注意力图的可信度。
AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

- 用对抗样本替换特征,避免传统掩码带来的偏差。
- 在多模态数据上验证,显著降低掩码引入的误差。
- 适合关注模型解释可信性的研究人员使用。
后处理注意力方法广泛用于解释深度神经网络,但其可信度难以可靠评估。现有评估通过按注意力排序遮蔽特征并测量性能下降,但该下降可能受遮蔽操作干扰:零值遮蔽会产生分布外伪影,插值遮蔽则可能保留残余预测信息。我们提出对抗性信息掩码(AIM),一种基于注意力引导的对抗特征替换框架,用于评估注意力图的可信度及遮蔽操作的可靠性。AIM 将选定特征替换为输入的对抗样本对应值,并比较互补遮蔽顺序下的性能下降。通过随机归因偏差和解释方法可信度排名的稳定性来评估可靠性。在图像、音频和脑电(EEG)任务上的实验表明,与零值和插值遮蔽相比,AIM 减少了遮蔽引起的偏差,同时揭示了符号化与非符号化注意力在不同模态间的差异。
原文摘要 · Abstract (English)
Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evaluations mask features according to saliency-induced feature ordering and measure performance degradation, but this degradation can be confounded by the masking operator: zero masking may create out-of-distribution artifacts, while interpolation-based masking may preserve residual predictive information. We propose Adversarial Information Masking (AIM), a saliency-guided adversarial feature replacement framework for evaluating both saliency-map faithfulness and masking-operator reliability. AIM replaces selected features with values from an adversarial counterpart of the input and compares degradation under complementary masking orders. We assess reliability using random-attribution bias and stability of explanation-method faithfulness rankings. Experiments on image, audio, and EEG tasks suggest that AIM reduces masking-induced bias compared with zero and interpolation-based masking, while revealing modality-dependent differences between signed and unsigned attributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。