将关键点描述子融入3D高斯泼溅,提升视觉定位精度
GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization
- 用XFeat提取关键点,融合进3DGS构建密集描述符
- 两阶段定位:先2D-3D匹配粗估姿态,再光度损失精修
- 在室内外数据集上优于NeRFMatch和PNeRFLoc
尽管已有多种视觉定位方法,如场景坐标回归和相机位姿回归,但这些方法常面临优化复杂或精度有限的问题。为解决上述挑战,我们探索了新型视图合成技术,特别是3D高斯泼溅(3DGS),其可紧凑编码三维几何与场景外观。我们提出一种两阶段流程,将轻量级XFeat特征提取器生成的密集且鲁棒的关键点描述子融入3DGS,显著提升室内与室外环境下的定位性能。粗略位姿通过3DGS表示与查询图像描述子之间的2D-3D对应关系直接获得。第二阶段则通过最小化基于渲染的光度扭曲损失,对初始位姿进行优化。在广泛使用的室内外数据集上的基准测试表明,该方法优于近期基于神经渲染的定位方法,如NeRFMatch和PNeRFLoc。
原文摘要 · Abstract (English)
Although various visual localization approaches exist, such as scene coordinate regression and camera pose regression, these methods often struggle with optimization complexity or limited accuracy. To address these challenges, we explore the use of novel view synthesis techniques, particularly 3D Gaussian Splatting (3DGS), which enables the compact encoding of both 3D geometry and scene appearance. We propose a two-stage procedure that integrates dense and robust keypoint descriptors from the lightweight XFeat feature extractor into 3DGS, enhancing performance in both indoor and outdoor environments. The coarse pose estimates are directly obtained via 2D-3D correspondences between the 3DGS representation and query image descriptors. In the second stage, the initial pose estimate is refined by minimizing the rendering-based photometric warp loss. Benchmarking on widely used indoor and outdoor datasets demonstrates improvements over recent neural rendering-based localization methods, such as NeRFMatch and PNeRFLoc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。