arXiv:2508.04987cs.CV2025-08TPAMI被引 9

解决视觉语言模型跨域适应中的模态差异问题,提升无监督域适应性能。

Unified modality separation: A vision-language framework for unsupervised domain adaptation

  • 分离视觉与语言特征中的共享与特定成分,统一处理。
  • 在多个数据集上实现最高9%的性能提升,计算效率提高9倍。
  • 适合研究跨模态对齐与无监督域适应的开发者使用。

无监督域适应(UDA)使在标注源域上训练的模型能够处理新的未标注目标域。近期预训练视觉语言模型(VLMs)通过利用语义信息,在零样本任务中表现出色。通过对齐视觉与文本嵌入,VLMs在弥合域间差距方面取得显著成效。然而,模态间固有的差异(即模态间隙)仍然存在。我们发现,存在模态间隙时直接进行UDA仅能迁移模态不变知识,导致目标域性能不佳。为此,我们提出一种统一的模态分离框架,同时处理模态特定与模态不变成分。训练阶段,从VLM特征中解耦不同模态组件,并以统一方式分别处理;测试时,自动确定模态自适应集成权重以最大化各组件协同效应。为评估实例级模态特性,设计模态差异度量,将样本分类为模态不变、模态特定及不确定三类。模态不变样本用于促进跨模态对齐,不确定样本则被标注以增强模型能力。基于提示调优技术,方法在计算效率提升9倍的同时,性能最高提升9%。大量实验与分析在多种骨干网络、基线、数据集和适应设置下验证了设计的有效性。

原文摘要 · Abstract (English)

Unsupervised domain adaptation (UDA) enables models trained on a labeled source domain to handle new unlabeled domains. Recently, pre-trained vision-language models (VLMs) have demonstrated promising zero-shot performance by leveraging semantic information to facilitate target tasks. By aligning vision and text embeddings, VLMs have shown notable success in bridging domain gaps. However, inherent differences naturally exist between modalities, which is known as modality gap. Our findings reveal that direct UDA with the presence of modality gap only transfers modality-invariant knowledge, leading to suboptimal target performance. To address this limitation, we propose a unified modality separation framework that accommodates both modality-specific and modality-invariant components. During training, different modality components are disentangled from VLM features then handled separately in a unified manner. At test time, modality-adaptive ensemble weights are automatically determined to maximize the synergy of different components. To evaluate instance-level modality characteristics, we design a modality discrepancy metric to categorize samples into modality-invariant, modality-specific, and uncertain ones. The modality-invariant samples are exploited to facilitate cross-modal alignment, while uncertain ones are annotated to enhance model capabilities. Building upon prompt tuning techniques, our methods achieve up to 9% performance gain with 9 times of computational efficiencies. Extensive experiments and analysis across various backbones, baselines, datasets and adaptation settings demonstrate the efficacy of our design.

域适应视觉语言模态分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。