arXiv:2506.11493cs.CV2025-06CVPR被引 12

通过视觉与文本嵌入的聚类特性,提升无监督域适应中的提示学习效果。

Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation

  • 利用源域与目标域视觉嵌入的关系,强化伪标签
  • 基于最优传输理论,使文本嵌入保持聚类结构
  • 适合关注提示学习与域适应的科研人员

近期利用CLIP等多模态预训练模型进行无监督域适应(UDA)的方法,通过丰富的语义知识和鲁棒的视觉表征,在跨域迁移中表现优异。尽管在多个基准上达到领先性能,但其提升主要依赖于基础伪标签(CLIP零样本预测)与自训练机制。然而,训练过程存在关键缺陷:目标域的视觉嵌入分布可能偏离预训练模型的分布,导致类别描述产生误导信号。本文提出新方法,通过挖掘视觉与文本嵌入间的几何结构——一种现有方法忽视的特性——来增强伪标签并促进目标提示学习。首先,我们直接利用源域提示的参考预测,基于源与目标视觉嵌入间的关系;随后发现预训练多模态模型中视觉与文本嵌入存在显著聚类行为。基于最优传输理论,我们将此发现转化为一种新策略,强制文本嵌入保持聚类性质,从而进一步提升目标域的对齐能力。实验与消融研究验证了该方法的有效性,表明其在表示质量与目标提示性能上均优于现有方法。

原文摘要 · Abstract (English)

Recent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-training on diverse image-text datasets. While these methods achieve state-of-the-art performance across benchmarks, much of the improvement stems from base pseudo-labels (CLIP zero-shot predictions) and self-training mechanisms. Thus, the training mechanism exhibits a key limitation wherein the visual embedding distribution in target domains can deviate from the visual embedding distribution in the pre-trained model, leading to misguided signals from class descriptions. This work introduces a fresh solution to reinforce these pseudo-labels and facilitate target-prompt learning, by exploiting the geometry of visual and text embeddings - an aspect that is overlooked by existing methods. We first propose to directly leverage the reference predictions (from source prompts) based on the relationship between source and target visual embeddings. We later show that there is a strong clustering behavior observed between visual and text embeddings in pre-trained multi-modal models. Building on optimal transport theory, we transform this insight into a novel strategy to enforce the clustering property in text embeddings, further enhancing the alignment in the target domain. Our experiments and ablation studies validate the effectiveness of the proposed approach, demonstrating superior performance and improved quality of target prompts in terms of representation.

无监督域适应提示学习多模态聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。