提出StyLIP与AD-CLIP,让视觉语言模型在跨域场景中更自适应。
In the Era of Prompt Learning with Vision-Language Models
- 用风格投影器分离图像内容与风格,学习领域无关的提示词
- 在多个域泛化基准上超越现有方法,无需微调即可跨域适配
- 适合需要少样本或无监督域适应的遥感、医疗等实际应用
大规模基础模型如CLIP虽具备强零样本泛化能力,但在面对领域偏移时表现受限。本文提出 extsc{StyLIP},一种面向域泛化(DG)的新型无领域提示学习策略。StyLIP通过风格投影器解耦CLIP视觉编码器中的视觉风格与内容,学习特定领域的提示词,并将其与内容特征结合。经对比学习训练后,该方法可实现跨域无缝适应,在多个DG基准上优于当前最优方法。此外,我们提出AD-CLIP用于无监督域适应(DA),利用冻结的CLIP视觉主干,通过图像风格与内容特征学习域不变提示词。借助嵌入空间中的熵最小化对齐不同域,即使仅有目标域样本,也能有效应对领域偏移。最后,我们展望了基于提示学习在遥感语义分割中发现新类别或罕见类别的未来工作,为复杂真实场景下更具适应性和泛化性的模型铺平道路。
原文摘要 · Abstract (English)
Large-scale foundation models like CLIP have shown strong zero-shot generalization but struggle with domain shifts, limiting their adaptability. In our work, we introduce \textsc{StyLIP}, a novel domain-agnostic prompt learning strategy for Domain Generalization (DG). StyLIP disentangles visual style and content in CLIP`s vision encoder by using style projectors to learn domain-specific prompt tokens and combining them with content features. Trained contrastively, this approach enables seamless adaptation across domains, outperforming state-of-the-art methods on multiple DG benchmarks. Additionally, we propose AD-CLIP for unsupervised domain adaptation (DA), leveraging CLIP`s frozen vision backbone to learn domain-invariant prompts through image style and content features. By aligning domains in embedding space with entropy minimization, AD-CLIP effectively handles domain shifts, even when only target domain samples are available. Lastly, we outline future work on class discovery using prompt learning for semantic segmentation in remote sensing, focusing on identifying novel or rare classes in unstructured environments. This paves the way for more adaptive and generalizable models in complex, real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。