arXiv:2411.15016cs.CVcs.RO2024-11中稿 · , code avaliable被引 30

融合4D雷达与摄像头,提升自动驾驶3D物体检测精度。

MSSF: A 4D Radar and Camera Fusion Framework With Multi-Stage Sampling for 3D Object Detection in Autonomous Driving

  • 分阶段采样融合网络,深度交互雷达点云与图像特征。
  • 在VoD和TJ4DRadSet上分别提升7.0%和4.0%的检测精度。
  • 适合追求低成本高鲁棒性的自动驾驶感知系统使用。

近年来兴起的4D毫米波雷达相比传统3D雷达具有更高分辨率并提供精确的仰角测量,但其点云仍稀疏且噪声大,难以满足自动驾驶需求。相机可捕捉丰富语义信息,因此4D雷达与相机融合能为自动驾驶系统提供低成本且鲁棒的感知方案。然而,现有雷达-相机融合方法尚未充分探索,性能与激光雷达方法仍有较大差距,主要因忽略特征模糊问题且未深入利用图像语义。为此,本文提出一种简单有效的多阶段采样融合(MSSF)网络。一方面设计可深度交互点云与图像特征的融合模块,支持通用单模态骨干网络即插即用,包含简单特征融合(SFF)与多尺度可变形特征融合(MSDFF)两种类型;另一方面提出语义引导头,在体素级进行前景-背景分割并重加权特征,缓解特征模糊问题。在View-of-Delft(VoD)与TJ4DRadset数据集上的大量实验验证了MSSF的有效性。显著地,相比当前最优方法,MSSF在VoD和TJ4DRadSet上分别提升7.0%和4.0%的3D平均精度,并在VoD上超越经典激光雷达方法。

原文摘要 · Abstract (English)

As one of the automotive sensors that have emerged in recent years, 4D millimeter-wave radar has a higher resolution than conventional 3D radar and provides precise elevation measurements. But its point clouds are still sparse and noisy, making it challenging to meet the requirements of autonomous driving. Camera, as another commonly used sensor, can capture rich semantic information. As a result, the fusion of 4D radar and camera can provide an affordable and robust perception solution for autonomous driving systems. However, previous radar-camera fusion methods have not yet been thoroughly investigated, resulting in a large performance gap compared to LiDAR-based methods. Specifically, they ignore the feature-blurring problem and do not deeply interact with image semantic information. To this end, we present a simple but effective multi-stage sampling fusion (MSSF) network based on 4D radar and camera. On the one hand, we design a fusion block that can deeply interact point cloud features with image features, and can be applied to commonly used single-modal backbones in a plug-and-play manner. The fusion block encompasses two types, namely, simple feature fusion (SFF) and multiscale deformable feature fusion (MSDFF). The SFF is easy to implement, while the MSDFF has stronger fusion abilities. On the other hand, we propose a semantic-guided head to perform foreground-background segmentation on voxels with voxel feature re-weighting, further alleviating the problem of feature blurring. Extensive experiments on the View-of-Delft (VoD) and TJ4DRadset datasets demonstrate the effectiveness of our MSSF. Notably, compared to state-of-the-art methods, MSSF achieves a 7.0% and 4.0% improvement in 3D mean average precision on the VoD and TJ4DRadSet datasets, respectively. It even surpasses classical LiDAR-based methods on the VoD dataset.

多传感器融合3D检测4D雷达自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。