arXiv:2608.26589cs.CV2026-08

用深度引导的投影对齐提升自动驾驶图像与点云配准精度

DPA-I2P: Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration in Autonomous Driving

论文配图:DPA-I2P: Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration in Autonomous Driving
图 1 · 摘自论文原文
  • 通过结构化几何感知方法融合深度与视觉信息,增强跨模态对齐
  • KITTI上相对最强基线,平移误差降45.0%,旋转误差降55.6%
  • 适合自动驾驶中复杂场景下的高精度定位任务

图像到点云配准旨在估计给定图像在三维场景点云中的相机位姿,是自动驾驶和大规模室外定位的基础任务。近期基于隐式对应学习的方法在端到端框架中提升了配准性能,实现了更精确的相机位姿估计。然而,由于图像与稀疏激光雷达点云之间固有的模态差异,可靠的跨模态对应学习仍具挑战。为此,本文提出深度引导的投影对齐方法(DPA-I2P)。不同于简单的深度或特征拼接,射线条件度量深度编码(RMDE)和投影一致视觉提升(PVL)以结构化、几何感知的方式利用深度与视觉线索。此外,跨模态查询剪枝(CQP)在早期精修阶段抑制不可靠查询,提升匹配稳定性。在KITTI和nuScenes上的实验表明该方法有效:在KITTI上,相比最强隐式基线,平均平移误差(RTE)降低45.0%,平均旋转误差(RRE)降低55.6%;在nuScenes上也优于所有对比基线,说明其在不同驾驶场景下具有更好泛化能力。

原文摘要 · Abstract (English)

Image-to-Point Cloud Registration aims to estimate the camera pose of a given image within a 3D scene point cloud, which is a fundamental task in autonomous driving and large-scale outdoor localization. Recent implicit correspondence learning methods have improved registration performance by learning cross-modal alignment in an end-to-end framework, leading to more accurate camera pose estimation. However, due to the inherent modality discrepancy between images and sparse LiDAR point clouds, reliable cross-modal correspondence learning remains challenging. To address this issue, we propose Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration (DPA-I2P). Unlike naive depth or feature concatenation, Ray-Conditioned Metric Depth Encoding (RMDE) and Projection-Consistent Vision Lifting (PVL) exploit depth and visual cues in a structured, geometry-aware manner. In addition, Cross-Modal Query Pruning (CQP) suppresses unreliable queries during early refinement to improve matching stability. Experiments on KITTI and nuScenes demonstrate the effectiveness of the proposed method. On KITTI, DPA-I2P reduces RTE and RRE by 45.0% and 55.6% over the strongest implicit baseline, respectively. On nuScenes, DPA-I2P also improves registration accuracy over the evaluated baselines, suggesting better transferability to different driving scenes.

点云配准自动驾驶多模态融合深度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。