arXiv:2509.21209cs.CVcs.LG2025-09

用置信预测方法让图像解释更可靠,用户可自定义解释精度。

Learning Conformal Explainers for Image Classifiers

  • 基于置信预测构建可调控精度的图像解释方法
  • FastSHAP在解释保真度和区域大小上均优于其他方法
  • 以超像素为单位的评估比像素级更有效,适合需要可信解释的研究

特征归因方法广泛用于解释图像分类结果,提供可直观可视化的特征级洞察。然而,这些解释的鲁棒性常不稳定,可能无法真实反映黑箱模型的推理过程。为此,我们提出一种基于置信预测的新方法,使用户可直接控制生成解释的保真度。该方法识别出一组足以维持模型预测的显著特征,无论被排除特征携带何种信息,且无需真实解释进行校准。提出四种一致性函数来量化解释与模型预测的一致程度。在六个图像数据集上对五种解释器进行实证评估,结果表明:FastSHAP在保真度和信息效率(以解释区域大小衡量)方面始终优于对比方法;此外,基于超像素的一致性度量比像素级更有效。

原文摘要 · Abstract (English)

Feature attribution methods are widely used for explaining image-based predictions, as they provide feature-level insights that can be intuitively visualized. However, such explanations often vary in their robustness and may fail to faithfully reflect the reasoning of the underlying black-box model. To address these limitations, we propose a novel conformal prediction-based approach that enables users to directly control the fidelity of the generated explanations. The method identifies a subset of salient features that is sufficient to preserve the model's prediction, regardless of the information carried by the excluded features, and without demanding access to ground-truth explanations for calibration. Four conformity functions are proposed to quantify the extent to which explanations conform to the model's predictions. The approach is empirically evaluated using five explainers across six image datasets. The empirical results demonstrate that FastSHAP consistently outperforms the competing methods in terms of both fidelity and informational efficiency, the latter measured by the size of the explanation regions. Furthermore, the results reveal that conformity measures based on super-pixels are more effective than their pixel-wise counterparts.

图像解释置信预测特征归因可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。