arXiv:2505.09915cs.CVcs.RO2025-05ICRA被引 10

用双目相机实现大规模户外3D高斯点云实时定位与建图

Large-Scale Gaussian Splatting SLAM

  • 基于多模态策略估计大视角变化下的初始位姿
  • 在EuRoc和KITTI数据集上优于现有神经与传统方法
  • 适合需要大场景高精度建图的自动驾驶与机器人应用

近期发展的神经辐射场(NeRF)和3D高斯点云(3DGS)在视觉SLAM中表现出色,但多数方法依赖RGBD传感器,仅适用于室内环境。大型室外场景下的重建鲁棒性仍待探索。本文提出基于双目相机的大规模3DGS视觉SLAM系统——LSG-SLAM。该系统采用多模态策略,在大视角变化下估计先验位姿;在跟踪阶段引入特征对齐的扭曲约束,缓解渲染损失中外观相似性带来的负面影响;为提升大场景可扩展性,提出连续高斯点云子地图,以有限内存应对无界场景。通过场景识别检测子地图间的回环,并利用渲染与特征扭曲损失优化闭环关键帧的相对位姿。全局位姿与高斯点优化后,结构精化模块进一步提升重建质量。在EuRoc与KITTI数据集上的大量实验表明,LSG-SLAM在性能上超越现有神经、3DGS及传统方法。

原文摘要 · Abstract (English)

The recently developed Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown encouraging and impressive results for visual SLAM. However, most representative methods require RGBD sensors and are only available for indoor environments. The robustness of reconstruction in large-scale outdoor scenarios remains unexplored. This paper introduces a large-scale 3DGS-based visual SLAM with stereo cameras, termed LSG-SLAM. The proposed LSG-SLAM employs a multi-modality strategy to estimate prior poses under large view changes. In tracking, we introduce feature-alignment warping constraints to alleviate the adverse effects of appearance similarity in rendering losses. For the scalability of large-scale scenarios, we introduce continuous Gaussian Splatting submaps to tackle unbounded scenes with limited memory. Loops are detected between GS submaps by place recognition and the relative pose between looped keyframes is optimized utilizing rendering and feature warping losses. After the global optimization of camera poses and Gaussian points, a structure refinement module enhances the reconstruction quality. With extensive evaluations on the EuRoc and KITTI datasets, LSG-SLAM achieves superior performance over existing Neural, 3DGS-based, and even traditional approaches. Project page: https://lsg-slam.github.io.

3D高斯视觉SLAM大场景建图双目相机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。