arXiv:2511.07974cs.AI2025-11CVPR被引 4

提出细粒度反事实解释框架,揭示模型误判的局部特征根源。

Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition

  • 基于相似性与贡献权重,非生成式生成对象与部件级解释。
  • 在细粒度分类任务中,显著提升对误判原因的解释粒度。
  • 适合需要精准定位误判原因的研究者,如医疗图像分析场景。

基于归因的解释方法虽能捕捉关键模式以增强视觉可解释性,但在细粒度任务中常缺乏足够细节,尤其在模型误判情况下,解释可能过于粗略。为解决此问题,我们提出一种细粒度反事实解释框架,实现对象级与部件级可解释性,回答两个核心问题:(1) 哪些细粒度特征导致模型误判?(2) 主导局部特征如何影响反事实调整?本方法通过量化正确分类与误分类样本在感兴趣区域内的相似性及组件贡献权重,非生成式地生成可解释反事实。此外,引入基于Shapley值的显著性分区模块,分离具有区域特异性相关性的特征。大量实验表明,该方法在捕捉更细粒度、直观有意义的区域方面优于现有细粒度方法。

原文摘要 · Abstract (English)

Attribution-based explanation techniques capture key patterns to enhance visual interpretability; however, these patterns often lack the granularity needed for insight in fine-grained tasks, particularly in cases of model misclassification, where explanations may be insufficiently detailed. To address this limitation, we propose a fine-grained counterfactual explanation framework that generates both object-level and part-level interpretability, addressing two fundamental questions: (1) which fine-grained features contribute to model misclassification, and (2) where dominant local features influence counterfactual adjustments. Our approach yields explainable counterfactuals in a non-generative manner by quantifying similarity and weighting component contributions within regions of interest between correctly classified and misclassified samples. Furthermore, we introduce a saliency partition module grounded in Shapley value contributions, isolating features with region-specific relevance. Extensive experiments demonstrate the superiority of our approach in capturing more granular, intuitively meaningful regions, surpassing fine-grained methods.

可解释性反事实解释细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。