arXiv:2503.19543cs.CV2025-03CVPR被引 8

无需重训或数据库,实现未知场景下的精准相机位姿估计

Scene-agnostic Pose Regression for Visual Localization

  • 提出无场景依赖的位姿回归新任务,双分支模型端到端预测
  • 在270个场景的360SPR数据集上,未知场景精度优于传统方法
  • 适用于无人重训的开放环境定位,适合机器人与自动驾驶

绝对位姿回归(APR)在未知环境中缺乏适应性,需重新训练;相对位姿回归(RPR)泛化较好但依赖大型图像检索库;视觉里程计(VO)在未见场景中表现良好,但存在轨迹累积误差。为此,本文提出新任务——无场景依赖位姿回归(SPR),可在无需重训或数据库的情况下灵活实现高精度位姿估计。为评估该任务,构建了大规模数据集360SPR,包含超过20万张全景图、360万张针孔图像及270个场景下三种传感器高度的相机位姿。进一步提出SPR-Mamba模型,采用双分支架构解决该问题。大量实验表明,本方法在360SPR与360Loc两个未知场景数据集上均持续优于APR、RPR和VO。相关代码与数据已公开。

原文摘要 · Abstract (English)

Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffers from accumulated error in open trajectories. To address this dilemma, we introduce a new task, Scene-agnostic Pose Regression (SPR), which can achieve accurate pose regression in a flexible way while eliminating the need for retraining or databases. To benchmark SPR, we created a large-scale dataset, 360SPR, with over 200K photorealistic panoramas, 3.6M pinhole images and camera poses in 270 scenes at three different sensor heights. Furthermore, a SPR-Mamba model is initially proposed to address SPR in a dual-branch manner. Extensive experiments and studies demonstrate the effectiveness of our SPR paradigm, dataset, and model. In the unknown scenes of both 360SPR and 360Loc datasets, our method consistently outperforms APR, RPR and VO. The dataset and code are available at https://junweizheng93.github.io/publications/SPR/SPR.html.

位姿估计无监督定位视觉定位三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。