arXiv:2510.20477cs.LG2025-10中稿 · IJCAI

通过双一致性机制提升视觉语言模型自训练的伪标签质量

Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models

  • 利用模型间与模型内一致性生成更可靠的伪标签
  • 在14个数据集上显著提升现有方法性能,尤其在小样本场景下
  • 无需预设阈值,适合各类视觉语言任务的微调

通过半监督学习(SSL)利用未标注数据,或通过微调预训练模型应对标签稀缺问题,是当前主流方法。近年来,将预训练视觉语言模型(VLMs)微调与半监督学习结合,形成新兴的半监督微调范式。然而,现有方法常因依赖预测一致性或预定义置信度阈值,导致模型偏差和超参数敏感性。为此,我们提出一种简单而高效的即插即用方法——双一致性引导自训练(Bi-CoG),通过同时利用模型间与模型内一致性,并结合误差感知的动态伪标签分配策略,生成高质量且低偏差的伪标签。理论分析与在14个数据集上的广泛实验表明,Bi-CoG能持续显著提升现有方法性能。

原文摘要 · Abstract (English)

Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for addressing label-scarce scenarios. Recently, growing attention has been given to combining fine-tuning of pre-trained vision-language models (VLMs) with SSL, forming the emerging paradigm of semi-supervised fine-tuning. However, existing methods often suffer from model bias and hyperparameter sensitivity, due to reliance on prediction consistency or pre-defined confidence thresholds. To address these limitations, we propose a simple yet effective plug-and-play methodology named $\underline{\textbf{Bi-Co}}$nsistency-$\underline{\textbf{G}}$uided Self-Training (Bi-CoG), which assigns high-quality and low-bias pseudo-labels, by simultaneously exploiting inter-model and intra-model consistency, along with an error-aware dynamic pseudo-label assignment strategy. Both theoretical analysis and extensive experiments over 14 datasets demonstrate the effectiveness of Bi-CoG, which consistently and significantly improves the performance of existing methods.

视觉语言模型自训练半监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。