arXiv:2509.15330cs.CV2025-09中稿 · TMLR 2026

用领域信息生成提示词,提升视觉语言模型的分布外泛化能力

CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization

  • 基于领域信息生成条件提示,增强图像与文本嵌入对齐
  • 在四个基准上实现显著更优的分布外性能,提升嵌入对齐度
  • 适合关注少样本、零样本跨域泛化的研究者

预训练视觉语言模型(如CLIP)在学习分布外(OOD)表征方面展现出巨大潜力。然而,基于提示的CLIP方法仍存在两大问题:一是文本描述不准确,导致精度和鲁棒性下降,尤其影响零样本场景;二是视觉-语言嵌入对齐不足,制约泛化性能。为此,本文提出条件领域提示学习(CoDoL),利用可获取的领域信息构建提示,促进更优的视觉-语言嵌入对齐,我们识别此为观察到的分布外泛化增益的重要因素。为进一步捕捉实例与领域双重特征,设计轻量级领域元网络(DMN),为各领域图像生成输入相关的提示词。在四个分布外基准(PACS、VLCS、OfficeHome、DigitDG)上的大量实验表明,所提CoDoL方法在四个基准上均实现了视觉-语言嵌入对齐的提升,验证了其有效性,并将此对齐改进视为推动分布外性能提升的关键贡献之一(非唯一原因)。

原文摘要 · Abstract (English)

Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potential in learning out-of-distribution (OOD) representations. Despite showing competitive performance, the prompt-based CLIP methods still suffer from: i) inaccurate text descriptions, which leads to degraded accuracy and robustness, and poses a challenge for zero-shot CLIP methods. ii) limited vision-language embedding alignment, which is one important factor affecting generalization performance. To tackle the above issues, this paper proposes a novel Conditional Domain prompt Learning (CoDoL) method, which utilizes readily-available domain information to form prompts and contributes to improved vision-language embedding alignment, which we identify as one factor underlying the observed OOD generalization gains. To capture both instance-specific and domain-specific information, we further propose a lightweight Domain Meta Network (DMN) to generate input-conditional tokens for images in each domain. Extensive experiments on four OOD benchmarks (PACS, VLCS, OfficeHome, and DigitDG) validate the effectiveness of our proposed CoDoL method in terms of empirically improves vision-language embedding alignment across four DG benchmarks, which we present as a contributing factor (rather than the sole cause) of the observed OOD gains.

视觉语言模型分布外泛化提示学习领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。