SimpleBEV通过优化相机与激光雷达融合,提升自动驾驶3D检测精度。
SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
- 在统一俯视图空间融合相机与激光雷达特征,改进编码器结构。
- 利用激光雷达校准相机深度估计,提升3D检测精度至77.6% NDS。
- 引入相机独立检测分支,增强训练阶段信息利用,适合自动驾驶系统开发。
越来越多研究将激光雷达与相机信息融合以提升自动驾驶系统的3D目标检测性能。近期一种简单而有效的融合框架在统一的俯视图(BEV)空间中实现了优异的检测效果。本文提出一种名为SimpleBEV的激光雷达-相机融合框架,用于高精度3D目标检测,该框架基于BEV融合范式,并分别改进了相机和激光雷达编码器。具体而言,采用级联网络进行基于相机的深度估计,并利用激光雷达点云获取的深度信息对结果进行校正。同时,引入一个仅使用相机-俯视图特征的辅助分支,在训练阶段充分挖掘相机信息。此外,通过融合多尺度稀疏卷积特征,改进激光雷达特征提取器。实验结果表明所提方法的有效性:在nuScenes数据集上达到77.6%的NDS准确率,显著优于现有方法。
原文摘要 · Abstract (English)
More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion framework has achieved an excellent detection performance, fusing the LiDAR and camera features in a unified bird's-eye-view (BEV) space. In this paper, we propose a LiDAR-camera fusion framework, named SimpleBEV, for accurate 3D object detection, which follows the BEV-based fusion framework and improves the camera and LiDAR encoders, respectively. Specifically, we perform the camera-based depth estimation using a cascade network and rectify the depth results with the depth information derived from the LiDAR points. Meanwhile, an auxiliary branch that implements the 3D object detection using only the camera-BEV features is introduced to exploit the camera information during the training phase. Besides, we improve the LiDAR feature extractor by fusing the multi-scaled sparse convolutional features. Experimental results demonstrate the effectiveness of our proposed method. Our method achieves 77.6\% NDS accuracy on the nuScenes dataset, showcasing superior performance in the 3D object detection track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。