arXiv:2505.07050cs.CV2025-05被引 1

利用深度图与图像的风格化融合,提升无监督领域泛化分割性能。

Depth-Sensitive Soft Suppression with RGB-D Inter-Modal Stylization Flow for Domain Generalization Semantic Segmentation

  • 通过跨模态风格化生成更鲁棒的深度图,检测敏感区域
  • 基于类别软抑制机制,增强不变特征学习,提升泛化能力
  • 首个融合RGB与深度信息的多类领域泛化分割框架,适合自动驾驶等场景

无监督域适应(UDA)需目标域数据以缩小域间差距,而域泛化(DG)无需目标数据。现有方法发现深度图有助于提升UDA表现,但忽略其因设备与环境因素导致的噪声和空洞问题,难以有效学习域不变特征。尽管高敏感区域抑制在学习不变特征方面表现良好,但现有方法无法直接应用于具有独特特性的深度图。为此,我们提出新颖框架DSSS(Depth-Sensitive Soft Suppression with RGB-D inter-modal stylization flow),聚焦于从深度图中学习域不变特征。具体地,设计了RGB-D跨模态风格化流,利用RGB信息作为风格化源生成样式化深度图以检测敏感区域;提出类别级软空间敏感性抑制机制,识别并强调含更多域不变信息的非敏感深度特征;同时引入RGB-D软对齐损失,确保样式化深度图仅部分对齐RGB特征,仍保留深度特异性。据我们所知,DSSS是首个在多类域泛化语义分割任务中整合RGB与深度信息的工作。在多个骨干网络上的大量实验表明,该框架实现了显著性能提升。

原文摘要 · Abstract (English)

Unsupervised Domain Adaptation (UDA) aims to align source and target domain distributions to close the domain gap, but still struggles with obtaining the target data. Fortunately, Domain Generalization (DG) excels without the need for any target data. Recent works expose that depth maps contribute to improved generalized performance in the UDA tasks, but they ignore the noise and holes in depth maps due to device and environmental factors, failing to sufficiently and effectively learn domain-invariant representation. Although high-sensitivity region suppression has shown promising results in learning domain-invariant features, existing methods cannot be directly applicable to depth maps due to their unique characteristics. Hence, we propose a novel framework, namely Depth-Sensitive Soft Suppression with RGB-D inter-modal stylization flow (DSSS), focusing on learning domain-invariant features from depth maps for the DG semantic segmentation. Specifically, we propose the RGB-D inter-modal stylization flow to generate stylized depth maps for sensitivity detection, cleverly utilizing RGB information as the stylization source. Then, a class-wise soft spatial sensitivity suppression is designed to identify and emphasize non-sensitive depth features that contain more domain-invariant information. Furthermore, an RGB-D soft alignment loss is proposed to ensure that the stylized depth maps only align part of the RGB features while still retaining the unique depth information. To our best knowledge, our DSSS framework is the first work to integrate RGB and Depth information in the multi-class DG semantic segmentation task. Extensive experiments over multiple backbone networks show that our framework achieves remarkable performance improvement.

域泛化语义分割深度图跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。