arXiv:2504.09202cs.CV2025-04被引 6

用扩散模型生成真实可信的反事实图像,解释分类器决策。

From Visual Explanations to Counterfactual Explanations with Latent Diffusion

  • 结合剪枝梯度攻击与潜空间扩散模型生成反事实图像。
  • 在ImageNet和CelebA-HQ上超越现有最佳方法,保持视觉真实性。
  • 适用于任意分类器,可揭示决策关键特征,适合可解释性研究者。

视觉反事实解释是理想的假设图像,能在高置信度下使分类器决策转向目标类别,同时保持视觉合理性和与原图的接近性。本文提出新方法解决当前主流工作中的两个关键挑战:一是确定区分目标类别与原始类别的关键反事实特征;二是无需依赖对抗鲁棒模型即可为非鲁棒分类器提供有效解释。我们的方法通过算法识别需修改的关键区域,再结合基于剪枝的对抗攻击与潜空间扩散模型生成逼真的反事实解释。该方法在ImageNet和CelebA-HQ数据集上的多项评估指标上均优于先前最先进结果。总体而言,本方法适用于任意分类器,强化了视觉与反事实解释之间的关联,实现语义有意义的更改,并向观察者呈现细微的反事实图像。

原文摘要 · Abstract (English)

Visual counterfactual explanations are ideal hypothetical images that change the decision-making of the classifier with high confidence toward the desired class while remaining visually plausible and close to the initial image. In this paper, we propose a new approach to tackle two key challenges in recent prominent works: i) determining which specific counterfactual features are crucial for distinguishing the "concept" of the target class from the original class, and ii) supplying valuable explanations for the non-robust classifier without relying on the support of an adversarially robust model. Our method identifies the essential region for modification through algorithms that provide visual explanations, and then our framework generates realistic counterfactual explanations by combining adversarial attacks based on pruning the adversarial gradient of the target classifier and the latent diffusion model. The proposed method outperforms previous state-of-the-art results on various evaluation criteria on ImageNet and CelebA-HQ datasets. In general, our method can be applied to arbitrary classifiers, highlight the strong association between visual and counterfactual explanations, make semantically meaningful changes from the target classifier, and provide observers with subtle counterfactual images.

反事实解释扩散模型可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。