多源多模态无监督域适应,提升自动驾驶3D目标检测泛化能力
MUSDA: Multi-source Multi-modality Unsupervised Domain Adaptive 3D Object Detection for Autonomous Driving

- 分层空间条件域分类器对齐相机与激光雷达特征
- 在Waymo、nuScenes、Lyft上均超越现有方法
- 适合跨场景自动驾驶系统快速部署
随着自动驾驶发展,大量多模态标注数据集已可用,为新环境下的3D目标检测提供了无需人工标注的域适应机会。然而传统域适应方法通常仅针对单一源域或单模态,难以应对多源多模态场景。本文提出一种新型多源多模态无监督域适应框架,利用多个带标签源域和一个无标签目标域,首先引入分层空间条件域分类器(HSC),对每对源-目标域在两个层次上联合对齐相机与激光雷达特征。为有效融合多源信息,构建各域对之间的原型图,并设计原型图加权(PGW)多源融合策略,聚合多个源检测头的预测结果。在Waymo、nuScenes和Lyft三个常用3D目标检测数据集上的实验表明,该框架能有效整合多模态与多源信息,持续优于当前最优方法。
原文摘要 · Abstract (English)
With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop domain-adaptive 3D object detectors for new environments without relying on labor-intensive manual annotations. However, traditional domain adaptation methods typically focus on a single source domain or a single modality, limiting their effectiveness in multi-source, multi-modality scenarios. In this paper, we propose a novel framework for multi-source, multi-modality unsupervised domain adaptation in 3D object detection for autonomous driving. Given multiple labeled source domains and one unlabeled target domain, our framework first introduces hierarchical spatially-conditioned (HSC) domain classifiers, which jointly align features from both camera and LiDAR modalities at two distinct levels for each source-target domain pair. To effectively leverage information from multiple source domains, we construct a prototype graph between each pair of domains. Based on this, we develop a prototype graph weighted (PGW) multi-source fusion strategy to aggregate predictions from multiple source detection heads. Experimental results on three widely used 3D object detection datasets - Waymo, nuScenes, and Lyft - demonstrate that our proposed framework effectively integrates information across both modalities and source domains, consistently outperforming state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。