提出快速精准的视觉反事实解释方法,提升模型可解释性。
MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual Explanations
- 通过自适应聚焦决策相关区域,实现局部精确生成。
- 推理速度比基线快30倍,保持高图像保真度。
- 无需训练,适用于多种视觉任务,实用性强。
视觉反事实解释旨在揭示能改变模型预测的最小语义修改,为深度神经网络提供因果且可解释的洞察。然而,现有基于扩散的反事实生成方法通常计算成本高、采样慢,且难以精确定位修改区域。为此,我们提出MaskDiME,一种简单、快速而有效的扩散框架,通过局部采样统一语义一致性和空间精度。该方法自适应地聚焦于决策相关区域,实现局部化且语义一致的反事实生成,同时保持高图像保真度。我们的训练无关框架在五个涵盖不同视觉领域的基准数据集上,推理速度比基线快30倍,性能相当或达到当前最优水平,为高效反事实解释提供了实用且通用的解决方案。
原文摘要 · Abstract (English)
Visual counterfactual explanations aim to reveal the minimal semantic modifications that can alter a model's prediction, providing causal and interpretable insights into deep neural networks. However, existing diffusion-based counterfactual generation methods are often computationally expensive, slow to sample, and imprecise in localizing the modified regions. To address these limitations, we propose MaskDiME, a simple, fast, yet effective diffusion framework that unifies semantic consistency and spatial precision through localized sampling. Our approach adaptively focuses on decision-relevant regions to achieve localized and semantically consistent counterfactual generation while preserving high image fidelity. Our training-free framework, MaskDiME, performs inference over 30x faster than the baseline and achieves comparable or state-of-the-art performance across five benchmark datasets spanning diverse visual domains, establishing a practical and generalizable solution for efficient counterfactual explanation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。