用图像生成法把稀疏激光雷达变密集图像,快速精准对齐相机与激光雷达。
Image-to-Point Cloud Registration Made Easy with Rectified Flow-based LiDAR Upsampling

- 将激光雷达视为成像设备,用条件修正流生成稠密强度图。
- 在R3LIVE上实现4.89°/1.63m均值误差,单次注册仅需0.68秒。
- 无需配对数据或真值姿态,可适配多种传感器组合。
图像到点云配准(I2P)对融合摄像头与激光雷达感知至关重要,但模态差异导致高精度与强泛化难以兼顾。本文提出一种简单高效的方法:将激光雷达视为成像传感器,从单次稀疏扫描中,利用条件修正流生成稠密激光雷达强度图像,再通过预训练特征匹配器与相机图像对齐,并用PnP-RANSAC估计6-自由度相对位姿。模型通过自监督图像补全任务预训练,仅需少量激光雷达数据微调(无需图像-点云配对数据或真值传感器姿态),即可适应多样化的激光雷达与相机配置。在R3LIVE数据集上的实验表明,该方法达到4.89°/1.63 m的平均误差,优于现有方法,且单次注册耗时约0.68秒。
原文摘要 · Abstract (English)
Image-to-Point Cloud Registration (I2P) is essential for integrating camera and LiDAR in perception and autonomous systems, yet the modality gap between images and point clouds makes it difficult to achieve both high accuracy and strong generalization. In this paper, we propose a simple yet effective I2P method that treats LiDAR as an imaging sensor: from a single sparse LiDAR scan, we generate a dense LiDAR intensity image using Conditional Rectified Flow, match it with a camera image using a pre-trained feature matcher, and estimate the 6-DoF relative pose via PnP-RANSAC. The proposed model is pre-trained through a self-supervised image completion task and fine-tuned on a small amount of LiDAR data (neither image-point cloud pairs nor ground-truth sensor poses are required), enabling it to scale to diverse LiDAR and camera configurations. Experiments on the R3LIVE dataset show that the proposed method achieves a mean error of 4.89° / 1.63 m, outperforming existing methods, while completing a single registration in approximately 0.68 s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。