arXiv:2510.19215cs.CV2025-10被引 2

用表面拟合增强雷达与相机融合,提升自动驾驶3D检测精度。

SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion

论文配图:SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
图 1 · 摘自论文原文
  • 通过二次曲面拟合显式建模物体几何,增强多模态交互。
  • 在TJ4DRadSet和VoD数据集上达到领先性能,显著改善深度估计。
  • 适合需要高精度3D检测的自动驾驶系统研发人员。

3D目标检测对自动驾驶至关重要。4D成像雷达成本低、探测距离远且能精确测速,但其点云稀疏、分辨率低,限制了物体几何表达并阻碍多模态融合。本文提出SFGFusion,一种由表面拟合引导的相机-4D成像雷达检测网络。通过从图像和雷达数据中估计物体的二次曲面参数,显式表面拟合模型增强了空间表征并促进跨模态交互,从而更可靠地预测细粒度稠密深度。该预测深度用于两个目的:1)在图像分支中指导图像特征从透视视图(PV)转换为统一鸟瞰图(BEV),提高空间映射精度;2)在表面伪点分支中生成稠密伪点云,缓解雷达点云稀疏问题。原始雷达点云在独立雷达分支中编码,并采用基于柱体的方法将其特征转换至BEV空间。最后,使用标准2D主干网络和检测头从BEV特征中预测目标类别和边界框。实验表明,SFGFusion有效融合相机与4D雷达特征,在TJ4DRadSet和view-of-delft(VoD)检测基准上均取得优异表现。

原文摘要 · Abstract (English)

3D object detection is essential for autonomous driving. As an emerging sensor, 4D imaging radar offers advantages as low cost, long-range detection, and accurate velocity measurement, making it highly suitable for object detection. However, its sparse point clouds and low resolution limit object geometric representation and hinder multi-modal fusion. In this study, we introduce SFGFusion, a novel camera-4D imaging radar detection network guided by surface fitting. By estimating quadratic surface parameters of objects from image and radar data, the explicit surface fitting model enhances spatial representation and cross-modal interaction, enabling more reliable prediction of fine-grained dense depth. The predicted depth serves two purposes: 1) in an image branch to guide the transformation of image features from perspective view (PV) to a unified bird's-eye view (BEV) for multi-modal fusion, improving spatial mapping accuracy; and 2) in a surface pseudo-point branch to generate dense pseudo-point cloud, mitigating the radar point sparsity. The original radar point cloud is also encoded in a separate radar branch. These two point cloud branches adopt a pillar-based method and subsequently transform the features into the BEV space. Finally, a standard 2D backbone and detection head are used to predict object labels and bounding boxes from BEV features. Experimental results show that SFGFusion effectively fuses camera and 4D radar features, achieving superior performance on the TJ4DRadSet and view-of-delft (VoD) object detection benchmarks.

3D检测多模态融合4D雷达自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。