arXiv:2505.13507cs.LGcs.CV2025-05被引 1

用视觉语言模型识别未知类别,提升跨域适应能力。

Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation

  • 通过可学习提示动态对齐源与目标域语义
  • 利用梯度范数差异区分已知与未知样本
  • 无需未知类标注,适合真实场景应用

开放集域适应(OSDA)需同时解决已知类别分布对齐与目标域未知类别的识别问题。现有方法常忽略模态间语义关系,且在未知样本检测中易积累误差。本文提出利用对比语言-图像预训练模型(CLIP)的两项创新:1)提示驱动的跨域对齐:基于域差异度量的可学习文本提示动态调整CLIP的文本编码器,实现源域与目标域间的语义一致性,无需未知类别监督;2)梯度感知的开放集分离:通过分析所学提示的梯度L2范数差异,已知与未知样本呈现统计上不同的梯度行为。在Office-Home数据集上的实验表明,该方法持续优于CLIP基线与标准基线。消融实验验证了梯度范数的关键作用。

原文摘要 · Abstract (English)

Open-Set Domain Adaptation (OSDA) confronts the dual challenge of aligning known-class distributions across domains while identifying target-domain-specific unknown categories. Current approaches often fail to leverage semantic relationships between modalities and struggle with error accumulation in unknown sample detection. We propose to harness Contrastive Language-Image Pretraining (CLIP) to address these limitations through two key innovations: 1) Prompt-driven cross-domain alignment: Learnable textual prompts conditioned on domain discrepancy metrics dynamically adapt CLIP's text encoder, enabling semantic consistency between source and target domains without explicit unknown-class supervision. 2) Gradient-aware open-set separation: A gradient analysis module quantifies domain shift by comparing the L2-norm of gradients from the learned prompts, where known/unknown samples exhibit statistically distinct gradient behaviors. Evaluations on Office-Home show that our method consistently outperforms CLIP baseline and standard baseline. Ablation studies confirm the gradient norm's critical role.

域适应视觉语言模型开放集学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。