arXiv:2411.05292cs.CVcs.AI2024-11被引 26

SimpleBEV通过优化相机与激光雷达融合,提升自动驾驶3D检测精度。

SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection

  • 在统一俯视图空间融合相机与激光雷达特征,改进编码器结构。
  • 利用激光雷达校准相机深度估计,提升3D检测精度至77.6% NDS。
  • 引入相机独立检测分支,增强训练阶段信息利用,适合自动驾驶系统开发。

越来越多研究将激光雷达与相机信息融合以提升自动驾驶系统的3D目标检测性能。近期一种简单而有效的融合框架在统一的俯视图(BEV)空间中实现了优异的检测效果。本文提出一种名为SimpleBEV的激光雷达-相机融合框架,用于高精度3D目标检测,该框架基于BEV融合范式,并分别改进了相机和激光雷达编码器。具体而言,采用级联网络进行基于相机的深度估计,并利用激光雷达点云获取的深度信息对结果进行校正。同时,引入一个仅使用相机-俯视图特征的辅助分支,在训练阶段充分挖掘相机信息。此外,通过融合多尺度稀疏卷积特征,改进激光雷达特征提取器。实验结果表明所提方法的有效性:在nuScenes数据集上达到77.6%的NDS准确率,显著优于现有方法。

原文摘要 · Abstract (English)

More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion framework has achieved an excellent detection performance, fusing the LiDAR and camera features in a unified bird's-eye-view (BEV) space. In this paper, we propose a LiDAR-camera fusion framework, named SimpleBEV, for accurate 3D object detection, which follows the BEV-based fusion framework and improves the camera and LiDAR encoders, respectively. Specifically, we perform the camera-based depth estimation using a cascade network and rectify the depth results with the depth information derived from the LiDAR points. Meanwhile, an auxiliary branch that implements the 3D object detection using only the camera-BEV features is introduced to exploit the camera information during the training phase. Besides, we improve the LiDAR feature extractor by fusing the multi-scaled sparse convolutional features. Experimental results demonstrate the effectiveness of our proposed method. Our method achieves 77.6\% NDS accuracy on the nuScenes dataset, showcasing superior performance in the 3D object detection track.

3D检测传感器融合自动驾驶BEV

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。