arXiv:2501.06831cs.CVcs.AI2025-01被引 13

通过分析模型内部特征,揭示图像分类决策的反事实与对比原因。

Towards Counterfactual and Contrastive Explainability and Transparency of DCNN Image Classifiers

  • 从模型内部筛选关键滤波器,生成可解释的对比与反事实说明
  • 在CUB-2011数据集上验证,能有效识别误分类的根本原因
  • 适合需要高透明度的医疗、金融等高风险场景使用

深度卷积神经网络(DCNN)的可解释性是重要研究方向,旨在揭示模型决策背后的依据,提升其在高风险环境中的可信度。本文提出一种新型方法,通过探测DCNN内部结构而非修改输入图像,生成可理解的反事实和对比解释。给定输入图像,该方法通过识别区分原预测类别与其他指定类别的重要滤波器,提供对比解释;同时,通过指定最小滤波器变化量,生成使模型输出发生改变的反事实解释。利用这些识别出的滤波器与概念,可揭示模型决策背后的对比与反事实依据,增强模型透明性。该方法可用于误分类分析,将特定输入的识别概念与类别特异性概念进行比较,验证模型判断的合理性。在Caltech-UCSD Birds (CUB) 2011数据集上与先进方法对比评估,验证了所生成解释的有效性。

原文摘要 · Abstract (English)

Explainability of deep convolutional neural networks (DCNNs) is an important research topic that tries to uncover the reasons behind a DCNN model's decisions and improve their understanding and reliability in high-risk environments. In this regard, we propose a novel method for generating interpretable counterfactual and contrastive explanations for DCNN models. The proposed method is model intrusive that probes the internal workings of a DCNN instead of altering the input image to generate explanations. Given an input image, we provide contrastive explanations by identifying the most important filters in the DCNN representing features and concepts that separate the model's decision between classifying the image to the original inferred class or some other specified alter class. On the other hand, we provide counterfactual explanations by specifying the minimal changes necessary in such filters so that a contrastive output is obtained. Using these identified filters and concepts, our method can provide contrastive and counterfactual reasons behind a model's decisions and makes the model more transparent. One of the interesting applications of this method is misclassification analysis, where we compare the identified concepts from a particular input image and compare them with class-specific concepts to establish the validity of the model's decisions. The proposed method is compared with state-of-the-art and evaluated on the Caltech-UCSD Birds (CUB) 2011 dataset to show the usefulness of the explanations provided.

可解释性反事实对比解释图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。