arXiv:2603.23159cs.CVcs.LG2026-03

用视觉语言模型指导图像标注,少标数据也能高精度。

Conformal Cross-Modal Active Learning

  • 用预训练视觉语言模型做教师,提供带语义的不确定性评分
  • 在多个数据集上用更少标注样本达到更高准确率
  • 适合需要减少人工标注成本的视觉模型训练场景

视觉领域的基础模型凭借强大的预训练表征和零样本能力,已显著提升图像识别性能,但其在数据高效学习方面的潜力仍未被充分挖掘。主动学习(AL)通过有策略地选择最信息量大的样本进行标注,以降低标注成本,但现有方法大多忽视了现代视觉-语言模型(VLMs)中丰富的多模态知识。本文提出一种新型主动学习框架——共形跨模态采样(CCMA),通过教师-学生架构连接视觉与语言模态。该框架利用预训练的视觉-语言模型作为教师,为视觉仅限的学生模型提供语义化的不确定性估计,并经共形校准以指导样本选择。通过融合多模态共形评分与多样性感知的选择策略,CCMA在多个基准测试中均展现出卓越的数据效率。实验表明,本方法持续优于当前最优的主动学习基线,明显优于仅依赖不确定度或多样性度量的方法。

原文摘要 · Abstract (English)

Foundation models for vision have transformed visual recognition with powerful pretrained representations and strong zero-shot capabilities, yet their potential for data-efficient learning remains largely untapped. Active Learning (AL) aims to minimize annotation costs by strategically selecting the most informative samples for labeling, but existing methods largely overlook the rich multimodal knowledge embedded in modern vision-language models (VLMs). We introduce Conformal Cross-Modal Acquisition (CCMA), a novel AL framework that bridges vision and language modalities through a teacher-student architecture. CCMA employs a pretrained VLM as a teacher to provide semantically grounded uncertainty estimates, conformally calibrated to guide sample selection for a vision-only student model. By integrating multimodal conformal scoring with diversity-aware selection strategies, CCMA achieves superior data efficiency across multiple benchmarks. Our approach consistently outperforms state-of-the-art AL baselines, demonstrating clear advantages over methods relying solely on uncertainty or diversity metrics.

主动学习跨模态视觉语言模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。