arXiv:2507.22075cs.LG2025-07

用原型与邻近一致性提升视觉语言模型无监督适配的伪标签质量

Prototype-Guided Pseudo-Labeling with Neighborhood-Aware Consistency for Unsupervised Adaptation

  • 基于类内紧凑性和类间分离性评估伪标签准确性
  • 通过邻居语义相似性动态修正伪标签,减少噪声
  • 自适应加权机制按正确性调节伪样本影响,适合复杂场景

在视觉语言模型(如CLIP)的无监督适配中,零样本预测生成的伪标签常含大量噪声,尤其在域偏移或视觉复杂场景下。传统依赖固定置信度阈值的过滤方法在完全无监督设置下不可靠。本文提出一种新型自适应伪标签框架,融合原型一致性(PICS)与邻域感知一致性(NALR),以提升CLIP的适配性能。PICS通过类内特征紧凑性和类间分离性评估伪标签准确性;NALR利用邻近样本间的语义相似性动态优化伪标签。此外,设计自适应加权机制,根据样本正确性估计调整其训练影响力。在11个基准数据集上的实验表明,该方法在无监督适配任务中达到当前最优性能,生成更准确伪标签的同时保持高效计算。

原文摘要 · Abstract (English)

In unsupervised adaptation for vision-language models such as CLIP, pseudo-labels derived from zero-shot predictions often exhibit significant noise, particularly under domain shifts or in visually complex scenarios. Conventional pseudo-label filtering approaches, which rely on fixed confidence thresholds, tend to be unreliable in fully unsupervised settings. In this work, we propose a novel adaptive pseudo-labeling framework that enhances CLIP's adaptation performance by integrating prototype consistency and neighborhood-based consistency. The proposed method comprises two key components: PICS, which assesses pseudo-label accuracy based on in-class feature compactness and cross-class feature separation; and NALR, which exploits semantic similarities among neighboring samples to refine pseudo-labels dynamically. Additionally, we introduce an adaptive weighting mechanism that adjusts the influence of pseudo-labeled samples during training according to their estimated correctness. Extensive experiments on 11 benchmark datasets demonstrate that our method achieves state-of-the-art performance in unsupervised adaptation scenarios, delivering more accurate pseudo-labels while maintaining computational efficiency.

无监督学习伪标签视觉语言模型域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。