用3D高斯点云实现遥感图像精确定位,兼顾速度与精度。
Hi^2-GSLoc: Dual-Hierarchical Gaussian-Specific Visual Relocalization for Remote Sensing

- 分层处理:先稀疏后稠密,利用高斯点特性逐步优化姿态估计。
- 在真实飞行数据中达到92.3%召回率,定位误差低于0.8米。
- 适合大规模遥感场景,支持实时计算,适合无人机导航应用。
视觉重定位通过查询图像估计相机的6自由度位姿,是遥感与无人机应用的基础。现有方法存在本质权衡:基于图像的检索与位姿回归精度不足,而基于结构的方法虽使用三维重建模型(如SfM),但计算复杂且难以扩展。在遥感场景中,这些挑战因大范围场景、高度变化和现有视觉先验的域差异更加显著。为此,本文采用3D高斯点阵(3DGS)作为新型场景表示,紧凑编码几何与外观信息。提出Hi²-GSLoc,一种双层次重定位框架,遵循稀疏到稠密、粗粒度到细粒度的范式,充分挖掘高斯原语中的语义与几何约束。为应对大规模遥感场景,引入分块高斯训练、GPU并行匹配与动态内存管理策略。方法分为两阶段:(1) 稀疏阶段,采用高斯特异性的一致性渲染感知采样与地标引导检测器,实现鲁棒准确的初始位姿估计;(2) 稠密阶段,通过粗到细的密集光栅化匹配迭代优化位姿,并加入可靠性验证。在仿真数据、公开数据集及真实飞行实验中全面评估,结果表明该方法在定位精度、召回率与计算效率上均具竞争力,能有效过滤不可靠位姿。验证了其在实际遥感应用中的有效性。
原文摘要 · Abstract (English)
Visual relocalization, which estimates the 6-degree-of-freedom (6-DoF) camera pose from query images, is fundamental to remote sensing and UAV applications. Existing methods face inherent trade-offs: image-based retrieval and pose regression approaches lack precision, while structure-based methods that register queries to Structure-from-Motion (SfM) models suffer from computational complexity and limited scalability. These challenges are particularly pronounced in remote sensing scenarios due to large-scale scenes, high altitude variations, and domain gaps of existing visual priors. To overcome these limitations, we leverage 3D Gaussian Splatting (3DGS) as a novel scene representation that compactly encodes both 3D geometry and appearance. We introduce $\mathrm{Hi}^2$-GSLoc, a dual-hierarchical relocalization framework that follows a sparse-to-dense and coarse-to-fine paradigm, fully exploiting the rich semantic information and geometric constraints inherent in Gaussian primitives. To handle large-scale remote sensing scenarios, we incorporate partitioned Gaussian training, GPU-accelerated parallel matching, and dynamic memory management strategies. Our approach consists of two stages: (1) a sparse stage featuring a Gaussian-specific consistent render-aware sampling strategy and landmark-guided detector for robust and accurate initial pose estimation, and (2) a dense stage that iteratively refines poses through coarse-to-fine dense rasterization matching while incorporating reliability verification. Through comprehensive evaluation on simulation data, public datasets, and real flight experiments, we demonstrate that our method delivers competitive localization accuracy, recall rate, and computational efficiency while effectively filtering unreliable pose estimates. The results confirm the effectiveness of our approach for practical remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。