融合4D雷达与相机数据,提升自动驾驶3D目标检测精度与鲁棒性
R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
- 通过全景深度融合增强绝对与相对深度估计
- 在无车辆位姿信息下仍保持稳定的时间融合性能
- 小目标检测依赖视觉先验,提升弱信号场景表现
4D雷达-相机感知配置在自动驾驶中日益重要。然而,现有融合4D雷达与相机数据的3D目标检测方法面临三大挑战:其一,绝对深度估计模块鲁棒性与精度不足,导致3D定位不准;其二,当本车位姿缺失或不准确时,时间融合模块性能显著下降甚至失效;其三,对于部分小目标,稀疏的雷达点云可能完全无法反射信号,此时检测只能依赖视觉单模态先验。为此,我们提出R4Det,通过全景深度融合模块提升深度估计质量,实现绝对与相对深度的相互增强。针对时间融合,设计无需依赖本车位姿的可变形门控时间融合模块。此外,构建实例引导动态精修模块,从2D实例引导中提取语义原型。实验表明,R4Det在TJ4DRadSet和VoD数据集上达到当前最优3D目标检测性能。代码与模型将开源于https://github.com/VDIGPKU/R4Det。
原文摘要 · Abstract (English)
4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fuse 4D Radar and camera data confront several challenges. First, their absolute depth estimation module is not robust and accurate enough, leading to inaccurate 3D localization. Second, the performance of their temporal fusion module will degrade dramatically or even fail when the ego vehicle's pose is missing or inaccurate. Third, for some small objects, the sparse radar point clouds may completely fail to reflect from their surfaces. In such cases, detection must rely solely on visual unimodal priors. To address these limitations, we propose R4Det, which enhances depth estimation quality via the Panoramic Depth Fusion module, enabling mutual reinforcement between absolute and relative depth. For temporal fusion, we design a Deformable Gated Temporal Fusion module that does not rely on the ego vehicle's pose. In addition, we built an Instance-Guided Dynamic Refinement module that extracts semantic prototypes from 2D instance guidance. Experiments show that R4Det achieves state-of-the-art 3D object detection results on the TJ4DRadSet and VoD datasets. The source code and models will be released at https://github.com/VDIGPKU/R4Det.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。