arXiv:2505.02056cs.CVcs.LG2025-05ICML

解决视觉语言模型伪标签不平衡问题,提升下游任务性能

Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin

  • 通过概念对齐和混淆感知的校准边界缓解伪标签偏差
  • 在6个数据集上相对最优方法提升6.29%准确率与平衡性
  • 适合需要高质量伪标签的自监督与半监督学习场景

将视觉语言模型(VLMs)适配到下游任务时使用伪标签已受到广泛关注。主要障碍在于VLM生成的伪标签往往存在分布不均,导致性能下降。尽管已有方法尝试多种策略应对,但不平衡的根本原因仍缺乏深入研究。为此,我们深入分析伪标签不平衡现象,识别出两个关键因素:概念错位与概念混淆。为缓解这两类问题,我们提出一种新框架,融合概念对齐与混淆感知的校准边界机制。其核心在于增强表现较差类别并促进各类别间预测的均衡性,从而缓解不平衡。在六个基准数据集及三种学习范式下的大量实验表明,该方法显著提升了伪标签的准确率与平衡性,相比当前最优方法实现6.29%的相对提升。代码已公开于https://anonymous.4open.science/r/CAP-C642/

原文摘要 · Abstract (English)

Adapting vision-language models (VLMs) to downstream tasks with pseudolabels has gained increasing attention. A major obstacle is that the pseudolabels generated by VLMs tend to be imbalanced, leading to inferior performance. While existing methods have explored various strategies to address this, the underlying causes of imbalance remain insufficiently investigated. To fill this gap, we delve into imbalanced pseudolabels and identify two primary contributing factors: concept mismatch and concept confusion. To mitigate these two issues, we propose a novel framework incorporating concept alignment and confusion-aware calibrated margin mechanisms. The core of our approach lies in enhancing underperforming classes and promoting balanced predictions across categories, thus mitigating imbalance. Extensive experiments on six benchmark datasets with three learning paradigms demonstrate that the proposed method effectively enhances the accuracy and balance of pseudolabels, achieving a relative improvement of 6.29% over the SoTA method. Our code is avaliable at https://anonymous.4open.science/r/CAP-C642/

视觉语言模型伪标签不平衡学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。