arXiv:2411.16407cs.CVcs.AI2024-11中稿 · British Machine Vi…

用视觉语言模型提升自动驾驶语义分割的无监督域适应能力

A Study on Unsupervised Domain Adaptation for Semantic Segmentation in the Era of Vision-Language Models

  • 用视觉语言预训练编码器替代原有模型,提升域适应性能
  • 在GTA5到Cityscapes任务上最高提升10.0% mIoU
  • 对未见领域泛化效果更好,适合自动驾驶真实场景迁移

尽管深度学习在计算机视觉领域取得进展,域偏移仍是主要挑战。自动驾驶语义分割面临天气变化、地理区域差异及合成数据使用带来的域偏移问题。无监督域适应(UDA)方法仅利用目标域无标签数据实现模型适配。现有方法普遍依赖ImageNet预训练模型,而视觉语言模型展现出更强泛化能力,可能助力域适应。我们发现,将现有UDA方法(如DACS)的编码器替换为视觉语言预训练编码器,可在GTA5到Cityscapes的域转移任务中带来最高达10.0% mIoU的性能提升。在三个未见数据集上,新编码器使泛化性能提升最高达13.7% mIoU。但并非所有UDA方法均可顺利适配新编码器,且域适应性能与泛化性能并不总能同步提升。我们在恶劣天气条件下的真实域转移任务中进一步验证了上述发现。

原文摘要 · Abstract (English)

Despite the recent progress in deep learning based computer vision, domain shifts are still one of the major challenges. Semantic segmentation for autonomous driving faces a wide range of domain shifts, e.g. caused by changing weather conditions, new geolocations and the frequent use of synthetic data in model training. Unsupervised domain adaptation (UDA) methods have emerged which adapt a model to a new target domain by only using unlabeled data of that domain. The variety of UDA methods is large but all of them use ImageNet pre-trained models. Recently, vision-language models have demonstrated strong generalization capabilities which may facilitate domain adaptation. We show that simply replacing the encoder of existing UDA methods like DACS by a vision-language pre-trained encoder can result in significant performance improvements of up to 10.0% mIoU on the GTA5-to-Cityscapes domain shift. For the generalization performance to unseen domains, the newly employed vision-language pre-trained encoder provides a gain of up to 13.7% mIoU across three unseen datasets. However, we find that not all UDA methods can be easily paired with the new encoder and that the UDA performance does not always likewise transfer into generalization performance. Finally, we perform our experiments on an adverse weather condition domain shift to further verify our findings on a pure real-to-real domain shift.

域适应语义分割视觉语言模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。