用3D高斯点云建图,直接估计相机位姿,提升低纹理环境定位精度
Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis

- 基于3D高斯溅射构建可微地图,直接回归6自由度位姿
- 融合多视图立体与激光雷达深度,几何损失提升重建精度
- 在类行星地貌中显著优于传统方法,适合火星车等机器人导航
在低纹理、感知混叠、光照恶劣且视角稀疏的类行星地形中,视觉定位面临严峻挑战。现有图像到图像或图像到地图匹配方法性能大幅下降。本文提出一种新型视觉重定位方法,摒弃传统特征匹配流程,直接基于3D高斯溅射(3DGS)构建的可微地图估计相机位姿。核心贡献在于引入一种几何感知训练策略,结合光度与几何损失,首次利用多视图立体(MVS)与激光雷达深度提供几何监督。联合优化使3DGS模型更贴合真实场景几何,提升光度与几何一致性,实现更鲁棒、精确的单图6-自由度位姿估计。在类行星环境采集的数据上进行大量实验,验证了该方法在复杂条件下的有效性,定位精度显著提升。代码已开源。
原文摘要 · Abstract (English)
Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation. In this work, we propose a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS). Our key contribution is a geometry-aware training strategy that combines photometric and geometric losses, where the geometric supervision is provided for the first time by combining multi-view stereo (MVS) and LiDAR depths. We show that this joint optimization produces a 3DGS model that better fits the underlying scene geometry, leading to improved photometric and geometric consistency and more robust, accurate single-image 6-DoF pose estimation. Extensive experiments on data acquired in planetary-analog environments validate the effectiveness of our approach, showing substantial gains in relocalization accuracy under challenging conditions. Code is available at https://github.com/DLR-RM/multimodal-gsplat-relocalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。