仅用一个视觉语言嵌入实现无目标数据的域适应,提升自动驾驶场景泛化能力。
Domain Adaptation with a Single Vision-Language Embedding
- 用单个视觉语言嵌入生成多种风格增强,通过优化低层特征仿射变换实现。
- 在零样本和单样本设置下,语义分割性能超越现有基线方法。
- 适用于真实驾驶场景中难以获取目标数据的情况,如恶劣天气条件。
域适应在计算机视觉中已被广泛研究,但在实际自动驾驶场景中,尤其在罕见或恶劣条件下,获取目标数据仍具挑战。本文提出一种新框架,仅依赖单一视觉-语言(VL)潜在嵌入进行域适应,无需完整目标数据。首先,基于对比语言-图像预训练模型(CLIP),提出提示/图像驱动的实例归一化(PIN)方法:通过优化源特征的仿射变换,利用单个目标域的VL嵌入挖掘多种视觉风格,实现特征增强。该嵌入可来自描述目标域的语言提示、部分优化的语言提示或单张未标注的目标图像。其次,证明这些挖掘出的风格(即增强)可用于零样本(即无目标数据)与单样本无监督域适应。在真实驾驶数据集Cityscapes和ACDC(恶劣条件)上的语义分割实验表明,该方法在实际的零样本与单样本设置下均优于相关基线。
原文摘要 · Abstract (English)
Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions. In this paper, we present a new framework for domain adaptation relying on a single Vision-Language (VL) latent embedding instead of full target data. First, leveraging a contrastive language-image pre-training model (CLIP), we propose prompt/photo-driven instance normalization (PIN). PIN is a feature augmentation method that mines multiple visual styles using a single target VL latent embedding, by optimizing affine transformations of low-level source features. The VL embedding can come from a language prompt describing the target domain, a partially optimized language prompt, or a single unlabeled target image. Second, we show that these mined styles (i.e., augmentations) can be used for zero-shot (i.e., target-free) and one-shot unsupervised domain adaptation. Experiments on semantic segmentation in real-world driving datasets, including Cityscapes and ACDC (adverse conditions), demonstrate the effectiveness of the proposed method, which outperforms relevant baselines in the practical zero-shot and one-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。