arXiv:2410.11211cs.CVcs.LG2024-10

融合相机与激光雷达特征,提升3D目标检测精度。

CVCP-Fusion: On Implicit Depth Estimation for 3D Bounding Box Prediction

  • 在鸟瞰图空间融合相机与激光雷达特征,保留语义密度。
  • 隐式深度估计在2D视图有效,但3D定位需显式空间信息。
  • 基于交叉视图注意力与中心点检测,支持实时处理。

结合激光雷达与摄像头数据已成为3D目标检测的常见方法。然而,以往方法在点级别融合两种输入流,丢失了来自摄像头特征的语义信息。本文提出跨视图中心点融合(CVCP-Fusion),一种先进的3D目标检测模型,在鸟瞰图(BEV)空间中融合摄像头与激光雷达特征,以保留摄像头流的语义密度,同时融入激光雷达流的空间数据。该架构借鉴了已有的交叉视图变换器与CenterPoint算法,采用并行运行其主干网络,实现高效计算,支持实时处理与应用。研究发现,尽管隐式深度估计在2D俯视图表示中已足够准确,但在3D世界视图中的精确边界框预测仍需显式的几何与空间信息。

原文摘要 · Abstract (English)

Combining LiDAR and Camera-view data has become a common approach for 3D Object Detection. However, previous approaches combine the two input streams at a point-level, throwing away semantic information derived from camera features. In this paper we propose Cross-View Center Point-Fusion, a state-of-the-art model to perform 3D object detection by combining camera and LiDAR-derived features in the BEV space to preserve semantic density from the camera stream while incorporating spacial data from the LiDAR stream. Our architecture utilizes aspects from previously established algorithms, Cross-View Transformers and CenterPoint, and runs their backbones in parallel, allowing efficient computation for real-time processing and application. In this paper we find that while an implicitly calculated depth-estimate may be sufficiently accurate in a 2D map-view representation, explicitly calculated geometric and spacial information is needed for precise bounding box prediction in the 3D world-view space.

3D检测多模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。