arXiv:2410.07475cs.CV2024-10CoRL被引 19

提出分阶段多模态融合框架,提升自动驾驶3D目标检测鲁棒性

Progressive Multi-Modal Fusion for Robust 3D Object Detection

论文配图:Progressive Multi-Modal Fusion for Robust 3D Object Detection
图 1 · 摘自论文原文
  • 在BEV和PV双视角下分层融合特征,保留高度与几何信息
  • 在nuScenes和Argoverse2上实现优于现有方法的检测精度
  • 支持单模态运行,对传感器故障有强容错能力

多传感器融合对自动驾驶中的精准3D目标检测至关重要,摄像头与激光雷达是最常用的传感器。然而,现有方法通常在单一视图(鸟瞰图BEV或透视图PV)中进行特征投影融合,导致高度和几何比例等互补信息丢失。为此,我们提出ProFusion3D,一种在中间特征和目标查询层面同时于BEV与PV进行渐进式融合的框架。该架构分层融合局部与全局特征,显著提升了3D目标检测的鲁棒性。此外,我们引入自监督掩码建模预训练策略,通过三个新目标提升多模态表征学习能力与数据效率。在nuScenes和Argoverse2数据集上的大量实验充分验证了ProFusion3D的有效性。更重要的是,当仅有一个模态可用时,ProFusion3D仍能保持优异性能。

原文摘要 · Abstract (English)

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspective View (PV), thus sacrificing complementary information such as height or geometric proportions. To address this limitation, we propose ProFusion3D, a progressive fusion framework that combines features in both BEV and PV at both intermediate and object query levels. Our architecture hierarchically fuses local and global features, enhancing the robustness of 3D object detection. Additionally, we introduce a self-supervised mask modeling pre-training strategy to improve multi-modal representation learning and data efficiency through three novel objectives. Extensive experiments on nuScenes and Argoverse2 datasets conclusively demonstrate the efficacy of ProFusion3D. Moreover, ProFusion3D is robust to sensor failure, demonstrating strong performance when only one modality is available.

3D检测多模态融合自动驾驶鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。