arXiv:2506.07327cs.CVcs.LG2025-06被引 1

提出新方法CASE,让模型解释更聚焦真实预测原因。

CASE: Contrastive Activation for Saliency Estimation

  • 设计对比机制,突出预测类别专属特征
  • 实验发现多数方法对不同类别输出相似解释
  • 适合需要可信解释的AI可解释性研究者

显著性方法广泛用于可视化模型关注的输入特征,但其视觉合理性可能掩盖关键缺陷。本文提出一种类别敏感性诊断测试:方法在相同输入下区分不同类别标签的能力。大量实验表明,许多常用显著性方法在不同类别下产生几乎相同的解释,质疑其可靠性。该类不敏感现象贯穿多种架构与数据集,暗示其为结构性缺陷而非模型特异性。受此启发,我们提出CASE——一种对比解释方法,能分离出对预测类别唯一有判别力的特征。通过所提诊断测试与基于扰动的保真度测试评估,CASE生成的解释更忠实且更具类别特异性。

原文摘要 · Abstract (English)

Saliency methods are widely used to visualize which input features are deemed relevant to a model's prediction. However, their visual plausibility can obscure critical limitations. In this work, we propose a diagnostic test for class sensitivity: a method's ability to distinguish between competing class labels on the same input. Through extensive experiments, we show that many widely used saliency methods produce nearly identical explanations regardless of the class label, calling into question their reliability. We find that class-insensitive behavior persists across architectures and datasets, suggesting the failure mode is structural rather than model-specific. Motivated by these findings, we introduce CASE, a contrastive explanation method that isolates features uniquely discriminative for the predicted class. We evaluate CASE using the proposed diagnostic and a perturbation-based fidelity test, and show that it produces faithful and more class-specific explanations than existing methods.

可解释性显著性图对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。