arXiv:2504.05254cs.CVcs.AI2025-04被引 1

用反事实图像解释模型为何不自信,提升可解释性。

Explaining Low Perception Model Competency with High-Competency Counterfactuals

  • 设计五种反事实图像生成方法,提升模型不确定性可视化
  • 三种方法(Reco/LGD/LNN)在六类低可信度场景中表现最优
  • 结合多模态大模型,反事实图像显著增强语言解释准确性

现有方法多关注如何解释图像分类模型的决策过程,但极少探讨模型为何缺乏信心。由于导致模型不确定的原因多样,不仅应识别其不确定性水平,还需解释原因。反事实图像可用于展示使分类结果改变所需做出的修改。本文提出五种生成高能力反事实图像的新方法:图像梯度下降(IGD)、特征梯度下降(FGD)、自编码器重构(Reco)、潜在空间梯度下降(LGD)和潜在最近邻(LNN)。我们在两个包含六类已知低模型能力成因的数据集上评估这些方法,发现Reco、LGD和LNN最具潜力。进一步测试表明,将反事实图像输入预训练多模态大语言模型(MLLM),能显著提升其对低感知能力原因的准确语言解释能力,证明反事实图像在解释模型低能力中的实用价值。

原文摘要 · Abstract (English)

There exist many methods to explain how an image classification model generates its decision, but very little work has explored methods to explain why a classifier might lack confidence in its prediction. As there are various reasons the classifier might lose confidence, it would be valuable for this model to not only indicate its level of uncertainty but also explain why it is uncertain. Counterfactual images have been used to visualize changes that could be made to an image to generate a different classification decision. In this work, we explore the use of counterfactuals to offer an explanation for low model competency--a generalized form of predictive uncertainty that measures confidence. Toward this end, we develop five novel methods to generate high-competency counterfactual images, namely Image Gradient Descent (IGD), Feature Gradient Descent (FGD), Autoencoder Reconstruction (Reco), Latent Gradient Descent (LGD), and Latent Nearest Neighbors (LNN). We evaluate these methods across two unique datasets containing images with six known causes for low model competency and find Reco, LGD, and LNN to be the most promising methods for counterfactual generation. We further evaluate how these three methods can be utilized by pre-trained Multimodal Large Language Models (MLLMs) to generate language explanations for low model competency. We find that the inclusion of a counterfactual image in the language model query greatly increases the ability of the model to generate an accurate explanation for the cause of low model competency, thus demonstrating the utility of counterfactual images in explaining low perception model competency.

可解释性反事实模型自信度多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。