用扩散模型生成图像回归的反事实解释,让预测结果更可懂。
Diffusion Counterfactuals for Image Regressors
- 基于像素和隐空间的扩散模型生成反事实图像
- 大预测值变化需大幅语义改变,反事实更难稀疏
- 适合研究模型决策机制或发现虚假关联的用户
反事实解释已成功用于提升各类黑箱模型的可解释性,尤其在图像领域受益于生成模型的进步。尽管其在分类任务中广泛应用,但在回归任务中的应用仍不充分。本文提出两种基于扩散生成模型的图像回归反事实解释方法:一种是直接在像素空间运行的去噪扩散概率模型,另一种是在隐空间操作的扩散自编码器。二者在CelebA-HQ和合成数据集上均生成了真实、语义合理且平滑的反事实图像,揭示了回归模型决策过程中的可解释洞察并暴露虚假相关性。研究发现,回归任务的反事实特征变化依赖于预测值区域;预测值显著变化需较大的语义变动,导致反事实难以稀疏化,相比分类任务更难生成。此外,像素空间反事实更稀疏,而隐空间反事实质量更高,支持更大语义变化。
原文摘要 · Abstract (English)
Counterfactual explanations have been successfully applied to create human interpretable explanations for various black-box models. They are handy for tasks in the image domain, where the quality of the explanations benefits from recent advances in generative models. Although counterfactual explanations have been widely applied to classification models, their application to regression tasks remains underexplored. We present two methods to create counterfactual explanations for image regression tasks using diffusion-based generative models to address challenges in sparsity and quality: 1) one based on a Denoising Diffusion Probabilistic Model that operates directly in pixel-space and 2) another based on a Diffusion Autoencoder operating in latent space. Both produce realistic, semantic, and smooth counterfactuals on CelebA-HQ and a synthetic data set, providing easily interpretable insights into the decision-making process of the regression model and reveal spurious correlations. We find that for regression counterfactuals, changes in features depend on the region of the predicted value. Large semantic changes are needed for significant changes in predicted values, making it harder to find sparse counterfactuals than with classifiers. Moreover, pixel space counterfactuals are more sparse while latent space counterfactuals are of higher quality and allow bigger semantic changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。