arXiv:2504.11930cs.CV2025-04CVPR

用扩散模型生成高质量伪标签,提升无监督提示学习的分类能力

Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning

  • 通过扩散模型生成高保真合成样本,构建辅助分类器增强判别力
  • 在5个公开数据集上,相比最优方法平均提升3.2%~5.7%准确率
  • 适合需要提升无监督视觉-语言对齐性能的研究者和工程师

利用大量未标注数据微调视觉语言模型(VLMs)近来受到广泛关注。然而,高质量伪标签的缺乏仍是关键挑战。当前伪标签策略常因语义与视觉信息不匹配,导致无监督提示学习(UPL)性能不佳。本文提出一种简单有效的方案——通过扩散模型增强判别丰富性(AiR),旨在更全面地表征类别,从而促进分类。具体而言,该方法包含一个伪标签生成模块,利用高保真合成样本构建辅助分类器,捕捉更丰富的视觉变化,将文本-图像对分类转化为更鲁棒的图像-图像对分类。同时,利用扩散生成样本的多样性增强提示学习,提供更多语义-视觉对齐信息。在五个公开基准(包括RESISC45和Flowers102)及三种学习范式(UL、SSL、TRZSL)上的实验表明,AiR显著且一致地超越现有最优无监督提示学习方法。

原文摘要 · Abstract (English)

Fine-tuning vision-language models (VLMs) with large amounts of unlabeled data has recently garnered significant interest. However, a key challenge remains the lack of high-quality pseudo-labeled data. Current pseudo-labeling strategies often struggle with mismatches between semantic and visual information, leading to sub-optimal performance of unsupervised prompt learning (UPL) methods. In this paper, we introduce a simple yet effective approach called \textbf{A}ugmenting D\textbf{i}scriminative \textbf{R}ichness via Diffusions (AiR), toward learning a richer discriminating way to represent the class comprehensively and thus facilitate classification. Specifically, our approach includes a pseudo-label generation module that leverages high-fidelity synthetic samples to create an auxiliary classifier, which captures richer visual variation, bridging text-image-pair classification to a more robust image-image-pair classification. Additionally, we exploit the diversity of diffusion-based synthetic samples to enhance prompt learning, providing greater information for semantic-visual alignment. Extensive experiments on five public benchmarks, including RESISC45 and Flowers102, and across three learning paradigms-UL, SSL, and TRZSL-demonstrate that AiR achieves substantial and consistent performance improvements over state-of-the-art unsupervised prompt learning methods.

无监督学习提示学习扩散模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。