arXiv:2605.16405cs.CV2026-05

用少量人工标注提升视觉语言模型引导的概念瓶颈模型性能

Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations

论文配图:Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations
图 1 · 摘自论文原文
  • 结合视觉语言模型与少量人工标注,通过高斯过程传播专家监督
  • 仅标注1%数据即实现更准确的概念预测与更好的可解释性
  • 适合需要高可解释性且标注成本高的实际应用场景

概念瓶颈模型(CBM)通过从输入中提取高层概念进行分类,确保决策可解释。但概念级标注稀缺,现有方法常依赖视觉语言模型(VLM)生成标注,导致概念质量下降。本文提出视觉-人类协同引导的CBM(VH-CBM),结合VLM与少量密集标注。在VLM嵌入空间中使用高斯过程捕捉目标域的全局信息,将专家标注推广至所有数据点。实验表明,即使仅标注1%数据,VH-CBM仍显著优于纯VLM引导的CBM,概念预测更准确、校准更好,并支持主动学习。

原文摘要 · Abstract (English)

Concept-bottleneck models (CBMs) are neural classifiers that compute predictions from high-level concepts extracted from the input. CBMs ensure stakeholders can understand the concepts -- and the predictions they entail -- by learning these from concept-level annotations, which are however seldom available. Recent CBM architectures work around this issue by obtaining annotations from Vision-Language Models (VLMs). While greatly broadening applicability, doing so can yield lower quality concepts and therefore less interpretable models. We strike for a middle ground by introducing Vision-plus-Human-guided CBM (VH-CBM), a hybrid approach that exploits both VLMs and a small amount of dense annotations. VH-CBM employs a Gaussian Process in the VLM's embedding space, which captures useful global information about the target domain, to propagate the expert's supervision to any target data point. Our empirical evaluation shows how VH-CBM predicts more accurate concepts than VLM-guided CBMs even when annotating as little as 1% of the data, while sporting better concept calibration and supporting active learning.

概念瓶颈可解释AI少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。