arXiv:2506.23975cs.CV2025-06

用实例相似性和概念相关性生成更简单可靠的图像分类解释

Toward Simple and Robust Contrastive Explanations for Image Classification by Leveraging Instance Similarity and Concept Relevance

  • 基于实例嵌入相似性和概念相关性提取解释
  • 高相关性概念产生更短、更简洁的解释
  • 在旋转、噪声等干扰下解释仍具一定鲁棒性

理解分类模型为何对某输入偏好某一类别而非另一类别,是对比解释的核心挑战。本文通过利用实例嵌入相似性和微调深度学习模型所使用的可人类理解概念的相关性,实现基于概念的对比解释。方法包括提取概念及其相关性得分,计算相似实例间的对比,并基于解释复杂度评估结果。实验验证了两个问题:(1) 解释复杂度是否随概念相关性变化;(2) 在旋转、噪声等图像增强下解释是否保持一致。结果表明,相关性越高,解释越短、越简洁;相关性越低,解释越长、越分散。此外,解释在不同增强下表现出不同程度的鲁棒性。这些发现为构建更可解释且鲁棒的人工智能系统提供了启示。

原文摘要 · Abstract (English)

Understanding why a classification model prefers one class over another for an input instance is the challenge of contrastive explanation. This work implements concept-based contrastive explanations for image classification by leveraging the similarity of instance embeddings and relevance of human-understandable concepts used by a fine-tuned deep learning model. Our approach extracts concepts with their relevance score, computes contrasts for similar instances, and evaluates the resulting contrastive explanations based on explanation complexity. Robustness is tested for different image augmentations. Two research questions are addressed: (1) whether explanation complexity varies across different relevance ranges, and (2) whether explanation complexity remains consistent under image augmentations such as rotation and noise. The results confirm that for our experiments higher concept relevance leads to shorter, less complex explanations, while lower relevance results in longer, more diffuse explanations. Additionally, explanations show varying degrees of robustness. The discussion of these findings offers insights into the potential of building more interpretable and robust AI systems.

对比解释可解释AI概念相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。