arXiv:2507.16224cs.CV2025-07被引 1

LDRFusion通过分阶段融合策略,提升复杂场景下3D目标检测精度。

LDRFusion: A LiDAR-Dominant multimodal refinement framework for 3D object detection

  • 先用激光雷达生成精准候选框,再引入伪点云细化难点目标
  • 在KITTI上多类别、多难度下均实现领先性能
  • 提出层次化残差编码模块,增强伪点云局部结构表达

现有激光雷达-相机融合方法在3D目标检测中表现优异。为缓解点云稀疏问题,以往方法通常通过深度补全构建空间伪点云作为辅助输入,并采用提议-精修框架生成检测结果。然而,伪点云会引入噪声,可能导致预测不准。鉴于各模态角色与可靠性差异,本文提出LDRFusion——一种新型激光雷达主导的两阶段精修框架。第一阶段仅依赖激光雷达生成精确定位的候选框;第二阶段引入伪点云以检测困难实例。两个阶段的实例级结果随后融合。为进一步提升伪点云的局部结构表征能力,提出层级伪点残差编码模块,通过特征与位置残差联合编码邻域集合。在KITTI数据集上的实验表明,本框架在多个类别和难度级别下均持续取得优异性能。

原文摘要 · Abstract (English)

Existing LiDAR-Camera fusion methods have achieved strong results in 3D object detection. To address the sparsity of point clouds, previous approaches typically construct spatial pseudo point clouds via depth completion as auxiliary input and adopts a proposal-refinement framework to generate detection results. However, introducing pseudo points inevitably brings noise, potentially resulting in inaccurate predictions. Considering the differing roles and reliability levels of each modality, we propose LDRFusion, a novel Lidar-dominant two-stage refinement framework for multi-sensor fusion. The first stage soley relies on LiDAR to produce accurately localized proposals, followed by a second stage where pseudo point clouds are incorporated to detect challenging instances. The instance-level results from both stages are subsequently merged. To further enhance the representation of local structures in pseudo point clouds, we present a hierarchical pseudo point residual encoding module, which encodes neighborhood sets using both feature and positional residuals. Experiments on the KITTI dataset demonstrate that our framework consistently achieves strong performance across multiple categories and difficulty levels.

3D检测多模态融合激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。