用雷达和激光雷达引导,提升摄像头3D目标检测精度。
RLG-TPV: Radar- and LiDAR-Guided Tri-Perspective View Fusion for Camera-Radar 3D Object Detection

- 融合雷达与激光雷达信息,指导图像特征的三维重建。
- 在nuScenes上达到0.4981 mAP,方向与速度误差降低超三成。
- 训练时用激光雷达,部署仅需摄像头和雷达,适合实际应用。
三视角视图(TPV)通过顶视、侧视和前视特征平面描述三维场景结构,但现有方法主要依赖摄像头,导致投影射线上的深度信息模糊。本文提出RLG-TPV,一种面向相机-雷达3D目标检测的多模态TPV框架,在表示构建过程中利用雷达与训练时的激光雷达提供互补几何引导。基于射线的可变形注意力机制,结合激光雷达监督的相机深度概率与雷达体素占用情况加权采样图像特征;雷达还进一步优化深度分布后再进行提升。由于传统雷达提供的俯仰信息有限,训练时使用激光雷达生成的类别占用目标监督侧视和前视平面,推理时移除对应检测头,部署仅需相机与雷达。时间聚合采用多普勒引导的时间融合,利用测量的雷达径向速度锚定运动场,并通过门控机制限制无运动支持区域的特征扭曲。此外,基于雷达散射截面(RCS)感知的雷达散射使雷达证据能在空间邻域内传播。在nuScenes验证集上,RLG-TPV实现0.4981 mAP和0.5959 NDS,相较于发布的CRN基线,姿态与速度误差分别降低31.9%和30.7%。消融实验表明,射线级几何引导是性能提升的关键因素。
原文摘要 · Abstract (English)
Tri-Perspective View (TPV) representations describe 3D scene structure through top, side, and front feature planes, but existing TPV lifting is primarily camera-based, leaving the depth of sampled image evidence ambiguous along projected camera rays. We propose RLG-TPV, a multimodal TPV framework for camera-radar 3D object detection in which radar and training-time LiDAR provide complementary geometric guidance during representation construction. A ray-guided deformable-attention lift weights sampled image features using LiDAR-supervised camera depth probabilities and radar frustum occupancy, while radar additionally refines the depth distribution before lifting. Because conventional radar provides limited elevation information, LiDAR-derived class-occupancy targets supervise the side and front planes during training; the corresponding heads are removed at inference, so deployment requires only cameras and radar. For temporal aggregation, Doppler-guided temporal fusion aligns past features using a motion field anchored by measured radar radial velocity, with gating that limits warping in regions without supported motion. An RCS-aware radar scatter further allows radar evidence to spread over spatial neighborhoods conditioned on radar cross section. On the nuScenes validation set, RLG-TPV achieves 0.4981 mAP and 0.5959 NDS, reducing orientation and velocity error by 31.9\% and 30.7\% relative to the published CRN baseline. Ablation studies show that ray-level geometric guidance is a major contributor to the final performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。