arXiv:2410.11505cs.CVcs.RO2024-10ICRA被引 10

用少量图像实现高精度视觉定位,基于高斯点云重建场景

LoGS: Visual Localization via Gaussian Splatting with Fewer Training Images

  • 用高斯点云表示场景,支持高质量新视角生成
  • 在少样本条件下仍达领先定位精度,四大数据集验证
  • 适合数据稀缺的机器人与AR应用

视觉定位旨在估计查询图像的6-自由度(6-DoF)相机位姿,是计算机视觉与机器人任务的基础。本文提出LoGS,一种基于3D高斯点云(Gaussian Splatting, GS)的视觉定位流程。该表示方法支持高质量的新视角合成。映射阶段先使用结构光恢复(SfM),再构建GS地图;定位阶段通过图像检索获取初始位姿,结合局部特征匹配与PnP求解器,最终通过基于合成分析的方法在GS地图上优化获得高精度位姿。在四个大规模数据集上的实验表明,该方法在相机位姿估计上达到当前最优性能,并在挑战性的少样本条件下表现出强鲁棒性。

原文摘要 · Abstract (English)

Visual localization involves estimating a query image's 6-DoF (degrees of freedom) camera pose, which is a fundamental component in various computer vision and robotic tasks. This paper presents LoGS, a vision-based localization pipeline utilizing the 3D Gaussian Splatting (GS) technique as scene representation. This novel representation allows high-quality novel view synthesis. During the mapping phase, structure-from-motion (SfM) is applied first, followed by the generation of a GS map. During localization, the initial position is obtained through image retrieval, local feature matching coupled with a PnP solver, and then a high-precision pose is achieved through the analysis-by-synthesis manner on the GS map. Experimental results on four large-scale datasets demonstrate the proposed approach's SoTA accuracy in estimating camera poses and robustness under challenging few-shot conditions.

视觉定位高斯点云少样本位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。