arXiv:2504.03438cs.CV2025-04CVPR被引 8

用4D雷达与视觉融合实现高精度3D目标检测,媲美激光雷达。

ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving

论文配图:ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving
图 1 · 摘自论文原文
  • 设计多尺度特征金字塔+双可变形交叉注意力融合模块。
  • 在VoD数据集上区域mAP达最新水平,整体性能接近激光雷达。
  • 低成本雷达替代方案,适合实际自动驾驶部署。

可靠的3D目标感知对自动驾驶至关重要。由于能在各种天气下工作,4D雷达近年来备受关注。然而,相比激光雷达,4D雷达提供的点云更稀疏。本文提出一种3D目标检测方法ZFusion,融合4D雷达与视觉模态。其核心FP-DDCA(特征金字塔-双可变形交叉注意力)融合器能有效互补稀疏雷达信息与密集视觉信息。具体而言,通过特征金字塔结构,将Transformer模块分层交互融合多尺度多模态特征,提升感知精度。此外,考虑到4D雷达的物理特性,引入深度-上下文-分割视图变换模块。由于4D雷达成本远低于激光雷达,ZFusion是激光雷达方案的有力替代。在典型交通场景如VoD数据集上,实验表明,该方法在合理推理速度下,区域内的平均精度(mAP)达到当前最优,全区域表现也优于基线方法,性能接近激光雷达,显著超越纯摄像头方法。

原文摘要 · Abstract (English)

Reliable 3D object perception is essential in autonomous driving. Owing to its sensing capabilities in all weather conditions, 4D radar has recently received much attention. However, compared to LiDAR, 4D radar provides much sparser point cloud. In this paper, we propose a 3D object detection method, termed ZFusion, which fuses 4D radar and vision modality. As the core of ZFusion, our proposed FP-DDCA (Feature Pyramid-Double Deformable Cross Attention) fuser complements the (sparse) radar information and (dense) vision information, effectively. Specifically, with a feature-pyramid structure, the FP-DDCA fuser packs Transformer blocks to interactively fuse multi-modal features at different scales, thus enhancing perception accuracy. In addition, we utilize the Depth-Context-Split view transformation module due to the physical properties of 4D radar. Considering that 4D radar has a much lower cost than LiDAR, ZFusion is an attractive alternative to LiDAR-based methods. In typical traffic scenarios like the VoD (View-of-Delft) dataset, experiments show that with reasonable inference speed, ZFusion achieved the state-of-the-art mAP (mean average precision) in the region of interest, while having competitive mAP in the entire area compared to the baseline methods, which demonstrates performance close to LiDAR and greatly outperforms those camera-only methods.

3D检测多模态融合4D雷达自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。