提出新任务OVDG-SS,让分割模型同时应对未知场景和未知类别。
Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
- 设计基于状态空间的文本图像关联优化机制S2-Corr,缓解域偏移导致的语义错配。
- 在自建基准上实现比现有方法更优的跨域泛化性能,尤其在合成到真实场景中提升显著。
- 适合自动驾驶、开放世界视觉等需要动态识别新类别与新环境的研究者。
语义分割中的领域泛化(DG-SS)旨在使模型在未见环境中保持鲁棒性。然而,传统方法受限于固定类别集合,难以适应开放世界场景。近年来,视觉语言模型(VLMs)推动了开放词汇语义分割(OV-SS),使模型能识别更广泛概念。但这些模型对域偏移仍敏感,在未见环境中表现下降,尤其在城市驾驶场景中问题突出。为此,我们提出开放词汇领域泛化语义分割(OVDG-SS),首次联合解决未知域与未知类别的挑战。构建了首个面向自动驾驶的OVDG-SS基准,涵盖从合成到真实、真实到真实的多种跨域泛化任务。实验发现,域偏移常扭曲预训练VLM中图文关联,阻碍OV-SS性能。为此,提出S2-Corr机制,通过状态空间驱动的图文关联精炼,有效缓解域偏移引起的失真,在分布变化下生成更一致的图文对应关系。大量实验证明,所提方法在跨域性能与效率上均优于现有方法。
原文摘要 · Abstract (English)
Domain Generalization in Semantic Segmentation (DG-SS) aims to enable segmentation models to perform robustly in unseen environments. However, conventional DG-SS methods are restricted to a fixed set of known categories, limiting their applicability in open-world scenarios. Recent progress in Vision-Language Models (VLMs) has advanced Open-Vocabulary Semantic Segmentation (OV-SS) by enabling models to recognize a broader range of concepts. Yet, these models remain sensitive to domain shifts and struggle to maintain robustness when deployed in unseen environments, a challenge that is particularly severe in urban-driving scenarios. To bridge this gap, we introduce Open-Vocabulary Domain Generalization in Semantic Segmentation (OVDG-SS), a new setting that jointly addresses unseen domains and unseen categories. We introduce the first benchmark for OVDG-SS in autonomous driving, addressing a previously unexplored problem and covering both synthetic-to-real and real-to-real generalization across diverse unseen domains and unseen categories. In OVDG-SS, we observe that domain shifts often distort text-image correlations in pre-trained VLMs, which hinders the performance of OV-SS models. To tackle this challenge, we propose S2-Corr, a state-space-driven text-image correlation refinement mechanism that mitigates domain-induced distortions and produces more consistent text-image correlations under distribution changes. Extensive experiments on our constructed benchmark demonstrate that the proposed method achieves superior cross-domain performance and efficiency compared to existing OV-SS approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。