融合4D雷达与相机数据,提升恶劣天气下3D目标检测精度
CVFusion: Cross-View Fusion of 4D Radar and Camera for 3D Object Detection
- 分两阶段融合:先用雷达引导生成候选框,再综合多视角特征优化
- 在VoD和TJ4DRadSet上分别提升9.10%和3.68%的mAP
- 适合自动驾驶中需高鲁棒性的多传感器融合场景
4D雷达因在恶劣天气下的强鲁棒性受到自动驾驶领域广泛关注。由于4D雷达点云稀疏且含噪,现有研究通常通过融合相机图像,在鸟瞰图(BEV)空间完成3D目标检测。然而,雷达潜力与融合机制仍未充分挖掘,制约了性能提升。本文提出一种跨视图两阶段融合网络CVFusion。第一阶段设计雷达引导的迭代式BEV融合模块(RGIter),生成高召回率的3D候选框;第二阶段对每个候选框聚合来自点云、图像和BEV空间的多源异构特征,实现精细化修正与高质量预测。在公开数据集上的大量实验表明,本方法显著超越现有最先进方法,在View-of-Delft(VoD)和TJ4DRadSet上分别取得9.10%和3.68%的mAP提升。代码将公开。
原文摘要 · Abstract (English)
4D radar has received significant attention in autonomous driving thanks to its robustness under adverse weathers. Due to the sparse points and noisy measurements of the 4D radar, most of the research finish the 3D object detection task by integrating images from camera and perform modality fusion in BEV space. However, the potential of the radar and the fusion mechanism is still largely unexplored, hindering the performance improvement. In this study, we propose a cross-view two-stage fusion network called CVFusion. In the first stage, we design a radar guided iterative (RGIter) BEV fusion module to generate high-recall 3D proposal boxes. In the second stage, we aggregate features from multiple heterogeneous views including points, image, and BEV for each proposal. These comprehensive instance level features greatly help refine the proposals and generate high-quality predictions. Extensive experiments on public datasets show that our method outperforms the previous state-of-the-art methods by a large margin, with 9.10% and 3.68% mAP improvements on View-of-Delft (VoD) and TJ4DRadSet, respectively. Our code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。