arXiv:2605.00480cs.CV2026-05中稿 · ICIP2026

用视觉语言模型生成粗粒度标签,降低细粒度分类的标注成本。

Leveraging Vision-Language Models as Weak Annotators in Active Learning

论文配图:Leveraging Vision-Language Models as Weak Annotators in Active Learning
图 1 · 摘自论文原文
  • 结合人类细粒度标注与模型粗粒度弱标注,动态分配标签。
  • 在CUB200和FGVC-Aircraft上,相同标注预算下性能优于现有方法。
  • 利用少量可信标签建模模型噪声,适合资源受限的图像分类任务。

主动学习旨在通过有选择地查询信息量大的样本,在有限标注预算下减少标注成本。本文研究如何利用视觉语言模型(VLMs)进一步降低对昂贵人工标注的依赖。实验发现,VLM在细粒度识别任务中,标签粒度越细,可靠性越低;而粗粒度标签则较准确。基于此特性,我们提出一种主动学习框架,通过实例级标签分配,融合细粒度人工标注与粗粒度VLM生成的弱标签。同时,利用少量可信全标注数据建模VLM标签中的系统性噪声。在CUB200和FGVC-Aircraft数据集上的实验表明,该框架在相同标注预算下持续优于现有主动学习方法。

原文摘要 · Abstract (English)

Active learning aims to reduce annotation cost by selectively querying informative samples for supervision under a limited labeling budget. In this work, we investigate how vision-language models (VLMs) can be leveraged to further reduce the reliance on costly human annotation within the active learning paradigm. To this end, we find that the reliability of VLMs varies significantly with label granularity in fine-grained recognition tasks: they perform poorly on fine-grained labels but can provide accurate coarse-grained labels. Leveraging this property, we propose an active learning framework that combines fine-grained human annotations with coarse-grained VLM-generated weak labels through instance-wise label assignment. We further model the systematic noise in VLM-generated labels using a small set of trusted full labels. Experiments on CUB200 and FGVC-Aircraft show that the proposed framework consistently outperforms existing active learning methods under the same annotation budget.

主动学习视觉语言模型弱监督细粒度分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。