用神经渲染实现静态激光雷达点云中实时精准摄像头定位。
CRISTAL: Real-time Camera Registration in Static LiDAR Scans using Neural Rendering
- 通过神经渲染生成合成视图,建立真实图像与点云的2D-3D对应关系。
- 在ScanNet++上实现无漂移、带量纲尺度的实时定位,优于现有SLAM方案。
- 适合需要高精度定位的机器人与扩展现实应用,无需标记物或闭环。
精准的相机定位对机器人和扩展现实(XR)至关重要,可保障可靠导航并实现虚拟与真实内容的对齐。现有视觉方法常受漂移、尺度模糊影响,且依赖标识物或回环检测。本文提出一种实时方法,在预先捕获的高精度彩色激光雷达点云中定位相机。通过从点云中渲染合成视图,建立实时帧与点云之间的2D-3D对应关系。采用神经渲染技术缩小合成图像与真实图像之间的域差距,减少遮挡与背景伪影,提升特征匹配效果。结果可在全局激光雷达坐标系中实现无漂移、具量纲尺度的相机跟踪。提出两种实时变体:Online Render and Match 和 Prebuild and Localize。在ScanNet++数据集上验证,性能优于现有SLAM系统。
原文摘要 · Abstract (English)
Accurate camera localization is crucial for robotics and Extended Reality (XR), enabling reliable navigation and alignment of virtual and real content. Existing visual methods often suffer from drift, scale ambiguity, and depend on fiducials or loop closure. This work introduces a real-time method for localizing a camera within a pre-captured, highly accurate colored LiDAR point cloud. By rendering synthetic views from this cloud, 2D-3D correspondences are established between live frames and the point cloud. A neural rendering technique narrows the domain gap between synthetic and real images, reducing occlusion and background artifacts to improve feature matching. The result is drift-free camera tracking with correct metric scale in the global LiDAR coordinate system. Two real-time variants are presented: Online Render and Match, and Prebuild and Localize. We demonstrate improved results on the ScanNet++ dataset and outperform existing SLAM pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。