arXiv:2509.18613cs.CV2025-09被引 5

融合4D雷达与相机,提升自动驾驶3D物体检测精度。

MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving

  • 分三层融合雷达点、场景和候选框信息,提升特征表达。
  • 在VoD数据集上达到接近激光雷达模型的检测性能。
  • 适合追求低成本高鲁棒性感知系统的自动驾驶研究者。

新兴的4D毫米波雷达可测量目标的距离、方位角、俯仰角和多普勒速度,在自动驾驶中因其成本低、鲁棒性强而备受关注。然而,其点云存在显著稀疏性和噪声,限制了单独使用于3D物体检测。现有4D雷达-相机融合方法大多沿用为激光雷达-相机设计的显式俯视图融合范式,忽视了雷达点云的稀疏与不完整特性,仅实现粗粒度场景级融合。为此,本文提出MLF-4DRCNet,一种两阶段的4D雷达与相机融合3D物体检测框架。模型融合点级、场景级和候选框级的多模态信息,包含三个关键模块:增强雷达点编码器(ERPE)、分层场景融合池化(HSFP)和候选框级融合增强(PLFE)。ERPE通过三重注意力体素特征编码器将雷达点云与2D图像实例对齐并密度化;HSFP利用可变形注意力动态融合多尺度体素特征与2D图像特征,并进行池化;PLFE通过融合图像特征优化区域建议,并与HSFP池化结果进一步融合。在View-of-Delft(VoD)和TJ4DRadSet数据集上的实验表明,MLF-4DRCNet达到当前最优性能,尤其在VoD数据集上表现接近激光雷达基线模型。

原文摘要 · Abstract (English)

The emerging 4D millimeter-wave radar, measuring the range, azimuth, elevation, and Doppler velocity of objects, is recognized for its cost-effectiveness and robustness in autonomous driving. Nevertheless, its point clouds exhibit significant sparsity and noise, restricting its standalone application in 3D object detection. Recent 4D radar-camera fusion methods have provided effective perception. Most existing approaches, however, adopt explicit Bird's-Eye-View fusion paradigms originally designed for LiDAR-camera fusion, neglecting radar's inherent drawbacks. Specifically, they overlook the sparse and incomplete geometry of radar point clouds and restrict fusion to coarse scene-level integration. To address these problems, we propose MLF-4DRCNet, a novel two-stage framework for 3D object detection via multi-level fusion of 4D radar and camera images. Our model incorporates the point-, scene-, and proposal-level multi-modal information, enabling comprehensive feature representation. It comprises three crucial components: the Enhanced Radar Point Encoder (ERPE) module, the Hierarchical Scene Fusion Pooling (HSFP) module, and the Proposal-Level Fusion Enhancement (PLFE) module. Operating at the point-level, ERPE densities radar point clouds with 2D image instances and encodes them into voxels via the proposed Triple-Attention Voxel Feature Encoder. HSFP dynamically integrates multi-scale voxel features with 2D image features using deformable attention to capture scene context and adopts pooling to the fused features. PLFE refines region proposals by fusing image features, and further integrates with the pooled features from HSFP. Experimental results on the View-of-Delft (VoD) and TJ4DRadSet datasets demonstrate that MLF-4DRCNet achieves the state-of-the-art performance. Notably, it attains performance comparable to LiDAR-based models on the VoD dataset.

自动驾驶4D雷达多模态融合3D检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。