用几何表示回归提升视觉定位精度与泛化能力
GRLoc: Geometric Representation Regression for Visual Localization
- 从图像直接回归3D几何表示,解耦旋转与平移预测
- 在7-Scenes和Cambridge Landmarks上达到顶尖性能
- 通过可微求解器恢复姿态,增强模型可解释性
绝对位姿回归(APR)已成为视觉定位的有力范式。然而,传统APR模型通常作为黑箱,直接从查询图像回归6自由度位姿,易陷入对训练视图的记忆而非理解三维场景几何结构。本文提出一种基于几何的替代方案:受新视角合成启发,将APR重构为其逆过程——直接从图像回归底层3D几何表示,称为几何表示回归(GRR)。模型显式预测世界坐标系下的两个解耦几何表示:(1) 射线方向图以估计相机旋转,(2) 对应点图以估计相机平移。最终位姿通过可微确定性求解器重建。这种解耦设计将几何先验引入网络,显著提升性能。在7-Scenes和Cambridge Landmarks数据集上实现当前最优结果,验证了建模逆渲染过程是更鲁棒的通用绝对位姿估计路径。
原文摘要 · Abstract (English)
Absolute Pose Regression (APR) has emerged as a compelling paradigm for visual localization. However, APR models typically operate as black boxes, directly regressing a 6-DoF pose from a query image, which can lead to memorizing training views rather than understanding 3D scene geometry. In this work, we propose a geometrically-grounded alternative. Inspired by novel view synthesis, which renders images from intermediate geometric representations, we reformulate APR as its inverse that regresses the underlying 3D representations directly from the image, and we name this paradigm Geometric Representation Regression (GRR). Our model explicitly predicts two disentangled geometric representations in the world coordinate system: (1) a raymap's directions to estimate camera rotation, and (2) a corresponding pointmap to estimate camera translation. The final camera pose is then recovered from these geometric components using a differentiable deterministic solver. This disentangled approach, which separates the learned visual-to-geometry mapping from the final pose calculation, introduces a strong geometric prior into the network. We find that the explicit decoupling of rotation and translation predictions measurably boosts performance. We demonstrate state-of-the-art performance on 7-Scenes and Cambridge Landmarks datasets, validating that modeling the inverse rendering process is a more robust path toward generalizable absolute pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。