提出新方法生成更可信的图像反事实解释,提升模型可解释性。
Towards Desiderata-Driven Design of Visual Counterfactual Explainers
- 融合多种机制构建平滑反事实生成算法
- 在合成与真实数据上验证了更高保真度与可理解性
- 适合关注模型透明性与可信解释的研究者
视觉反事实解释器(VCE)是一种直观且有前景的方法,用于提升图像分类器的透明性。它们通过揭示模型响应最强烈的特定数据变换,补充特征重要性等其他解释方式。本文指出,现有VCE过度关注样本质量或变化最小性,忽视了更全面的解释期望,如保真度、可理解性和充分性。为此,我们探索新的反事实生成机制,并研究其如何满足这些目标。将这些机制整合为一种新型的‘平滑反事实探索者’(SCE)算法,并在合成与真实数据上通过系统评估验证了其有效性。
原文摘要 · Abstract (English)
Visual counterfactual explainers (VCEs) are a straightforward and promising approach to enhancing the transparency of image classifiers. VCEs complement other types of explanations, such as feature attribution, by revealing the specific data transformations to which a machine learning model responds most strongly. In this paper, we argue that existing VCEs focus too narrowly on optimizing sample quality or change minimality; they fail to consider the more holistic desiderata for an explanation, such as fidelity, understandability, and sufficiency. To address this shortcoming, we explore new mechanisms for counterfactual generation and investigate how they can help fulfill these desiderata. We combine these mechanisms into a novel 'smooth counterfactual explorer' (SCE) algorithm and demonstrate its effectiveness through systematic evaluations on synthetic and real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。