arXiv:2410.22187cs.CV2024-10中稿 · WACV 2025被引 18

用主动学习选关键样本,提升视觉语言模型零样本性能

Active Learning for Vision-Language Models

  • 通过校准熵并结合自不确定性和邻域不确定性,更可靠地选样本
  • 在多个图像分类数据集上优于现有主动学习方法,显著提升模型表现
  • 适合想用少量标注数据优化预训练视觉语言模型的研究者

预训练的视觉语言模型(如CLIP)在众多下游计算机视觉任务中展现出出色的零样本性能。然而,其性能仍与在下游数据集上监督训练的深度模型存在显著差距。为缩小这一差距,我们提出一种新型主动学习(AL)框架,通过从无标签数据中选择少量信息量高的样本进行标注,来增强VLM的零样本分类能力。该方法首先校准VLM的预测熵,再结合自不确定性和邻域感知不确定性,计算出可靠的不确定性度量用于样本选择。大量实验表明,所提方法在多个图像分类数据集上优于现有AL方法,并显著提升VLM的零样本性能。

原文摘要 · Abstract (English)

Pre-trained vision-language models (VLMs) like CLIP have demonstrated impressive zero-shot performance on a wide range of downstream computer vision tasks. However, there still exists a considerable performance gap between these models and a supervised deep model trained on a downstream dataset. To bridge this gap, we propose a novel active learning (AL) framework that enhances the zero-shot classification performance of VLMs by selecting only a few informative samples from the unlabeled data for annotation during training. To achieve this, our approach first calibrates the predicted entropy of VLMs and then utilizes a combination of self-uncertainty and neighbor-aware uncertainty to calculate a reliable uncertainty measure for active sample selection. Our extensive experiments show that the proposed approach outperforms existing AL approaches on several image classification datasets, and significantly enhances the zero-shot performance of VLMs.

主动学习视觉语言模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。