用高斯点云合成视角,提升稀疏图像下的定位精度
Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting
- 用特征高斯点云表示场景,自适应合成更匹配查询视角的虚拟图像
- 混合使用渲染特征与原始图像特征,在粗/精阶段分别优化匹配效果
- 在室内外数据集上实现新最优性能,尤其适用于低密度图像场景
视觉重定位是3D计算机视觉中的基础任务,旨在当相机再次访问已知场景时估计其位姿。尽管基于点的分层重定位方法表现出良好的可扩展性和效率,但常受限于数据库图像稀疏和特征匹配能力弱。本文提出SplatHLoc,一种新颖的分层视觉重定位框架,采用特征高斯点云作为场景表示。为解决数据库图像稀疏问题,提出自适应视点检索方法,合成与查询更对齐的虚拟候选视图,从而提升初始位姿估计精度。针对特征匹配,观察到高斯渲染特征与直接从图像提取的特征在两阶段匹配中表现不同:前者在粗阶段更优,后者在细阶段更有效。因此引入混合特征匹配策略,实现更准确高效的位姿估计。在室内外数据集上的大量实验表明,SplatHLoc显著增强了视觉重定位的鲁棒性,达到新的最先进水平。
原文摘要 · Abstract (English)
Visual relocalization is a fundamental task in the field of 3D computer vision, estimating a camera's pose when it revisits a previously known scene. While point-based hierarchical relocalization methods have shown strong scalability and efficiency, they are often limited by sparse image observations and weak feature matching. In this work, we propose SplatHLoc, a novel hierarchical visual relocalization framework that uses Feature Gaussian Splatting as the scene representation. To address the sparsity of database images, we propose an adaptive viewpoint retrieval method that synthesizes virtual candidates with viewpoints more closely aligned with the query, thereby improving the accuracy of initial pose estimation. For feature matching, we observe that Gaussian-rendered features and those extracted directly from images exhibit different strengths across the two-stage matching process: the former performs better in the coarse stage, while the latter proves more effective in the fine stage. Therefore, we introduce a hybrid feature matching strategy, enabling more accurate and efficient pose estimation. Extensive experiments on both indoor and outdoor datasets show that SplatHLoc enhances the robustness of visual relocalization, setting a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。