arXiv:2511.00067cs.LGcs.AI2025-11中稿 · ICASSP 2026被引 1

无需领域标签,通过隐式领域聚类提升视觉语言模型泛化能力

Latent Domain Prompt Learning for Vision-Language Models

  • 用图像特征自动发现隐式领域并聚类,构建可迁移的领域表征
  • 在4个基准上超越基线模型,显著提升跨领域鲁棒性
  • 适合需要无标签域适应的现实应用,如跨场景图像理解

域泛化(DG)的目标是使模型对域偏移具备鲁棒性,这对视觉语言模型(VLMs)在真实场景中的部署至关重要。然而,现有方法大多依赖于可能不可用或模糊的域标签。本文研究无显式域标签条件下的域泛化问题。核心思想是将未见目标域表示为训练数据中自动发现的隐式域的组合,从而实现模型在域间的自适应知识迁移。为此,我们在图像特征上进行隐式域聚类,并根据输入图像与每个隐式域的相似度融合特定域的文本特征。在四个基准上的实验表明,该策略持续优于基于VLM的基线模型,并为提升域偏移下的鲁棒性提供了新见解。

原文摘要 · Abstract (English)

The objective of domain generalization (DG) is to enable models to be robust against domain shift. DG is crucial for deploying vision-language models (VLMs) in real-world applications, yet most existing methods rely on domain labels that may not be available and often ambiguous. We instead study the DG setting where models must generalize well without access to explicit domain labels. Our key idea is to represent an unseen target domain as a combination of latent domains automatically discovered from training data, enabling the model to adaptively transfer knowledge across domains. To realize this, we perform latent domain clustering on image features and fuse domain-specific text features based on the similarity between the input image and each latent domain. Experiments on four benchmarks show that this strategy yields consistent gains over VLM-based baselines and provides new insights into improving robustness under domain shift.

域泛化视觉语言模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。