用雷达解决相机后向投影的深度模糊问题,提升3D检测精度。
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection based View Transformation
- 结合相机图像与雷达数据,通过后向投影增强深度分辨能力。
- 在nuScenes数据集上达到62.4% NDS和54.0% mAP,性能领先。
- 适合自动驾驶中需要高精度3D感知的场景使用。
基于相机与雷达融合的鸟瞰图(BEV)3D目标检测方法因传感器互补性和成本优势受到关注。以往采用前向投影的方法面临BEV特征稀疏的问题,而使用后向投影的方法则忽略深度模糊,导致误检。本文提出一种新型相机-雷达融合模型CRAB(Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation),利用雷达信息缓解后向投影中的深度模糊。在视图变换过程中,CRAB将透视图图像上下文特征聚合至BEV查询中,通过融合图像提供的稠密但不可靠深度分布与雷达提供的稀疏但精确深度信息,提升同射线查询间的深度区分度。此外,引入含雷达上下文信息的特征图进行空间交叉注意力,增强对3D场景的理解。在nuScenes公开数据集上,该方法在后向投影类相机-雷达融合方法中达到当前最优性能,3D目标检测NDS为62.4%,mAP为54.0%。
原文摘要 · Abstract (English)
Recently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4\% NDS and 54.0\% mAP in 3D object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。