arXiv:2504.06800cs.CV2025-04被引 4

用图像生成技术精准扰动高相关区域,更真实评估模型解释性。

A Meaningful Perturbation Metric for Evaluating Explainability Methods

  • 用图像修复模型只修改高相关像素,保持图像自然性。
  • 新方法在多种模型和方法上产生更可信的排序结果。
  • 排序结果与人类偏好相关性显著高于现有方法,适合可解释性研究者。

深度神经网络虽表现优异,但决策过程不透明,阻碍其广泛应用。为此,归因方法被提出以分配输入各部分的相关性值。然而不同方法产生的相关性图差异大,需标准化评估指标。现有评估多通过扰动实现,即修改高/低相关区域观察预测变化。本文提出新方法:利用图像生成模型仅对输入图像的高相关像素进行修复式扰动,改变模型预测同时保持图像保真度。相比现有方法常产生分布外修改、导致不可靠结果的问题,本方法能生成更合理的排名。大量实验表明,该方法在多种模型和归因方法下均有效,且其生成的排序与人类偏好相关性显著更高,凸显其在提升DNN可解释性方面的潜力。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) have demonstrated remarkable success, yet their wide adoption is often hindered by their opaque decision-making. To address this, attribution methods have been proposed to assign relevance values to each part of the input. However, different methods often produce entirely different relevance maps, necessitating the development of standardized metrics to evaluate them. Typically, such evaluation is performed through perturbation, wherein high- or low-relevance regions of the input image are manipulated to examine the change in prediction. In this work, we introduce a novel approach, which harnesses image generation models to perform targeted perturbation. Specifically, we focus on inpainting only the high-relevance pixels of an input image to modify the model's predictions while preserving image fidelity. This is in contrast to existing approaches, which often produce out-of-distribution modifications, leading to unreliable results. Through extensive experiments, we demonstrate the effectiveness of our approach in generating meaningful rankings across a wide range of models and attribution methods. Crucially, we establish that the ranking produced by our metric exhibits significantly higher correlation with human preferences compared to existing approaches, underscoring its potential for enhancing interpretability in DNNs.

可解释性归因方法图像生成评价指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。