arXiv:2505.07209cs.CV2025-05CVPR被引 15

通过解耦最优传输模型,精准捕捉图像局部区域与概念的细粒度关联。

Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models

  • 将概念预测建模为局部图像块与概念间的运输问题,实现细粒度对齐。
  • 在多个任务上达到最佳性能,提升分类与概念预测准确性。
  • 适合关注可解释性、细粒度分析及抗数据偏差的视觉模型研究者。

概念瓶颈模型(CBMs)通过探索输入图像与输出预测之间的中间概念空间,使决策过程透明化。现有CBMs仅学习整体图像与概念间的粗粒度关系,忽略局部信息,导致两个主要问题:一是产生虚假的视觉-概念关联,降低模型可靠性;二是虽能解释各概念对预测的重要性,却难以定位具体贡献的视觉区域。为此,本文提出解耦最优传输概念瓶颈模型(DOT-CBM),以探索局部图像块与概念间的细粒度视觉-概念关系。具体地,将概念预测过程建模为图像块与概念间的运输问题,实现显式的细粒度特征对齐;同时在模态内引入正交投影损失,增强局部特征解耦。为进一步缓解数据统计偏见引发的捷径问题,利用视觉显著图和概念标签统计作为运输先验。因此,DOT-CBM可生成反演热图,提供更可靠的的概念预测,并实现更准确的类别预测。大量实验表明,该方法在图像分类、局部部件检测及分布外泛化等多个任务上均达到当前最优性能。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image information, leading to two main drawbacks: i) they often produce spurious visual-concept relations, hence decreasing model reliability; and ii) though CBMs could explain the importance of every concept to the final prediction, it is still challenging to tell which visual region produces the prediction. To solve these problems, this paper proposes a Disentangled Optimal Transport CBM (DOT-CBM) framework to explore fine-grained visual-concept relations between local image patches and concepts. Specifically, we model the concept prediction process as a transportation problem between the patches and concepts, thereby achieving explicit fine-grained feature alignment. We also incorporate orthogonal projection losses within the modality to enhance local feature disentanglement. To further address the shortcut issues caused by statistical biases in the data, we utilize the visual saliency map and concept label statistics as transportation priors. Thus, DOT-CBM can visualize inversion heatmaps, provide more reliable concept predictions, and produce more accurate class predictions. Comprehensive experiments demonstrate that our proposed DOT-CBM achieves SOTA performance on several tasks, including image classification, local part detection and out-of-distribution generalization.

可解释性细粒度分析最优传输概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。