让概念瓶颈模型精准定位图像中的视觉概念。
Locality-aware Concept Bottleneck Model
- 为每个概念设计原型,通过基础模型学习其典型局部特征。
- 在不降低分类性能前提下,显著提升概念定位精度。
- 适合需要可解释性和空间定位的视觉理解任务。
概念瓶颈模型(CBM)通过人类可理解的视觉概念进行预测,具有内在可解释性。由于密集标注概念成本高昂,近期方法利用基础模型自动识别图像中存在的概念。然而,这类无标签CBM常无法准确定位概念所在区域,预测时会关注无关视觉区域。为此,本文提出局部感知概念瓶颈模型(LCBM),利用基础模型的丰富信息,并引入原型学习机制,确保概念的空间定位准确。具体而言,为每个概念分配一个原型,该原型被训练为表示该概念的典型图像特征。通过鼓励原型编码相似局部区域,并借助基础模型保证原型与对应概念的相关性,从而引导模型学习正确的概念预测区域。实验表明,LCBM能有效识别图像中出现的概念,在保持相近分类性能的同时显著改善定位效果。
原文摘要 · Abstract (English)
Concept bottleneck models (CBMs) are inherently interpretable models that make predictions based on human-understandable visual cues, referred to as concepts. As obtaining dense concept annotations with human labeling is demanding and costly, recent approaches utilize foundation models to determine the concepts existing in the images. However, such label-free CBMs often fail to localize concepts in relevant regions, attending to visually unrelated regions when predicting concept presence. To this end, we propose a framework, coined Locality-aware Concept Bottleneck Model (LCBM), which utilizes rich information from foundation models and adopts prototype learning to ensure accurate spatial localization of the concepts. Specifically, we assign one prototype to each concept, promoted to represent a prototypical image feature of that concept. These prototypes are learned by encouraging them to encode similar local regions, leveraging foundation models to assure the relevance of each prototype to its associated concept. Then we use the prototypes to facilitate the learning process of identifying the proper local region from which each concept should be predicted. Experimental results demonstrate that LCBM effectively identifies present concepts in the images and exhibits improved localization while maintaining comparable classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。