针对自动驾驶多模态3D全景分割的域适应问题,提出首个专门框架PanDA。
PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

- 设计非对称多模态增强,模拟传感器退化以提升鲁棒性。
- 引入双专家伪标签优化模块,提升伪标签完整性和可靠性。
- 在光照、天气、地点等多类域偏移下超越现有最佳方法。
本文首次研究了自动驾驶场景中多模态3D全景分割(mm-3DPS)的无监督域适应(UDA)。现有方法依赖强跨模态互补性(激光雷达与图像),在单模态退化(如恶劣天气或弱光)时表现脆弱。传统伪标签仅保留高置信区域,导致掩码碎片化,严重影响全景分割。为此,提出PanDA,首个专为mm-3DPS设计的UDA框架。通过不对称多模态增强,选择性丢弃区域以模拟域偏移,提升表示学习鲁棒性;设计双专家伪标签精炼模块,从2D与3D模态中提取领域不变先验,增强伪标签的完整性与可靠性。在涵盖时间、天气、位置及传感器变化的多种域偏移设置下,显著优于当前最先进的3D语义分割UDA基线。
原文摘要 · Abstract (English)
This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strategy, which is widely used in UDA to generate supervision for unlabeled target data, combined with an mm-3DPS backbone. However, existing supervised mm-3DPS methods rely heavily on strong cross-modal complementarity between LiDAR and RGB inputs, making them fragile under domain shifts where one modality degrades (e.g., poor lighting or adverse weather). Moreover, conventional pseudo-labeling typically retains only high-confidence regions, leading to fragmented masks and incomplete object supervision, which are issues particularly detrimental to panoptic segmentation. To address these challenges, we propose PanDA, the first UDA framework specifically designed for multimodal 3D panoptic segmentation. To improve robustness against single-sensor degradation, we introduce an asymmetric multimodal augmentation that selectively drops regions to simulate domain shifts and improve robust representation learning. To enhance pseudo-label completeness and reliability, we further develop a dual-expert pseudo-label refinement module that extracts domain-invariant priors from both 2D and 3D modalities. Extensive experiments across diverse domain shifts, spanning time, weather, location, and sensor variations, significantly surpass state-of-the-art UDA baselines for 3D semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。