arXiv:2512.16164cs.CVcs.AI2025-12

通过双分支对齐提升视觉语言模型在无监督域适应中的分类性能

C-DGPA: Class-Centric Dual-Alignment Generative Prompt Adaptation

  • 设计双分支架构,同时优化边缘分布与条件分布对齐
  • 在OfficeHome等三个数据集上达到新最佳性能
  • 适合需要跨域泛化能力的视觉语言模型研究者

无监督域自适应将标注源域的知识迁移至未标注目标域。直接在下游任务中使用视觉语言模型并结合提示调优,面临缓解领域差异的重大挑战。现有提示调优策略主要关注边缘分布对齐,忽视条件分布差异,导致类别原型错位和语义判别力下降。为此,本文提出类中心双对齐生成提示适配方法(C-DGPA)。该方法通过新颖的双分支架构,协同优化边缘分布与条件分布对齐。边缘分布对齐分支采用动态对抗训练框架以弥合边缘分布差异;条件分布对齐分支引入类映射机制(CMM),通过标准化语义提示理解并防止对源域过度依赖,从而对齐条件分布差异。该双对齐策略通过协同优化,有效将领域知识融入提示学习,确保领域不变且语义可区分的表示。在OfficeHome、Office31和VisDA-2017上的大量实验验证了其优越性,所有基准均取得新的最优结果。

原文摘要 · Abstract (English)

Unsupervised Domain Adaptation transfers knowledge from a labeled source domain to an unlabeled target domain. Directly deploying Vision-Language Models (VLMs) with prompt tuning in downstream UDA tasks faces the signifi cant challenge of mitigating domain discrepancies. Existing prompt-tuning strategies primarily align marginal distribu tion, but neglect conditional distribution discrepancies, lead ing to critical issues such as class prototype misalignment and degraded semantic discriminability. To address these lim itations, the work proposes C-DGPA: Class-Centric Dual Alignment Generative Prompt Adaptation. C-DGPA syner gistically optimizes marginal distribution alignment and con ditional distribution alignment through a novel dual-branch architecture. The marginal distribution alignment branch em ploys a dynamic adversarial training framework to bridge marginal distribution discrepancies. Simultaneously, the con ditional distribution alignment branch introduces a Class Mapping Mechanism (CMM) to align conditional distribu tion discrepancies by standardizing semantic prompt under standing and preventing source domain over-reliance. This dual alignment strategy effectively integrates domain knowl edge into prompt learning via synergistic optimization, ensur ing domain-invariant and semantically discriminative repre sentations. Extensive experiments on OfficeHome, Office31, and VisDA-2017 validate the superiority of C-DGPA. It achieves new state-of-the-art results on all benchmarks.

域适应提示调优视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。