arXiv:2509.16567cs.CVcs.AI2025-09NeurIPS被引 3

无需训练,用理论指导编辑生成可解释的视觉反事实。

V-CECE: Visual Counterfactual Explanations via Conceptual Edits

  • 基于理论最优编辑步骤,逐步修改图像内容。
  • 在CNN、ViT和LVLM上均实现人类级解释效果。
  • 无需访问分类器内部,适合黑箱模型可解释性研究。

现有黑箱反事实生成框架忽视编辑内容的语义,过度依赖训练。本文提出一种新式即插即用的黑箱反事实生成框架,通过理论保证的最优编辑步骤生成人类可理解的反事实解释,且无需训练。该框架利用预训练图像编辑扩散模型,不需访问分类器内部结构,实现可解释的反事实生成过程。实验中,我们通过全面的人类评估,揭示了人类推理与神经模型行为之间的解释差距,验证了方法在卷积神经网络(CNN)、视觉变换器(ViT)和大视觉语言模型(LVLM)上的有效性。

原文摘要 · Abstract (English)

Recent black-box counterfactual generation frameworks fail to take into account the semantic content of the proposed edits, while relying heavily on training to guide the generation process. We propose a novel, plug-and-play black-box counterfactual generation framework, which suggests step-by-step edits based on theoretical guarantees of optimal edits to produce human-level counterfactual explanations with zero training. Our framework utilizes a pre-trained image editing diffusion model, and operates without access to the internals of the classifier, leading to an explainable counterfactual generation process. Throughout our experimentation, we showcase the explanatory gap between human reasoning and neural model behavior by utilizing both Convolutional Neural Network (CNN), Vision Transformer (ViT) and Large Vision Language Model (LVLM) classifiers, substantiated through a comprehensive human evaluation.

反事实解释扩散模型可解释AI视觉编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。