arXiv:2608.10938cs.CV2026-08中稿 · IROS 2026

用3D高斯点云统一估计相机6自由度位姿,精度与泛化性兼备。

GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting

论文配图:GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting
图 1 · 摘自论文原文
  • 先通过几何检索粗估位姿,再用高斯点云逐层优化精修
  • 在多个室内室外数据集上达到当前最优精度
  • 适合需要高精度位姿估计的机器人与AR应用

尽管视觉定位领域已从场景坐标回归发展到直接相机位姿回归,但实现鲁棒泛化与高精度仍具挑战。本文提出基于3D高斯点云的相机位姿估计方法GS-CPE,采用由粗到精的框架:首先在3DGS场景表示上通过检索引导的几何方法粗估位姿,随后在多尺度优化框架中,通过感知可见性的掩码RGB扭曲目标进行位姿精修,并结合自适应重渲染。在7Scenes、Cambridge Landmarks、FAST-LIVO2等室内室外基准以及自建数据集上的大量实验表明,该方法在精度和泛化能力上均显著优于现有方法,达到当前最优水平。

原文摘要 · Abstract (English)

Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generalization and high accuracy remain challenging. This study introduces GS-CPE (Gaussian Splatting based Camera Pose Estimation), a coarse-to-fine framework for 6-DoF camera pose estimation that unifies geometry-based coarse pose estimation with robust 3D Gaussian Splatting (3DGS) warping based pose refinement. GS-CPE first estimates a coarse pose via retrieval-guided geometric pose estimation on a 3DGS scene representation, then refines it by minimizing a visibility aware masked RGB warping objective in a multi-scale optimization framework, with adaptive re-rendering. Extensive experiments on indoor and outdoor benchmarks including 7Scenes, Cambridge Landmarks, FAST-LIVO2 datasets, and a custom dataset demonstrate state-of-the-art performance, consistently outperforming in both accuracy and generalization.

位姿估计3D高斯视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。