arXiv:2603.07937cs.CV2026-03被引 1

无需提前建图,实时重建即可精准定位野外场景。

$L^3$:Scene-agnostic Visual Localization in the Wild

  • 直接在线重建3D结构,跳过传统预处理步骤。
  • 在稀疏场景下表现优于主流方法,定位精度接近顶尖水平。
  • 适合无预建地图的移动设备定位,如无人机、机器人巡检。

标准视觉定位方法通常需要对场景进行离线预处理以获取三维结构信息,这带来了额外的计算成本和存储开销。能否在无需任何离线预处理的情况下实现野外场景的视觉定位?本文提出一种新型免地图视觉定位框架 $L^3$,利用前馈式3D重建网络的在线推理能力,通过对RGB图像直接进行在线3D重建,并基于2D-3D对应关系完成两阶段度量尺度恢复与位姿精修,实现了无需预先构建或存储任何场景表示的高精度定位。大量实验表明,$L^3$ 在多个基准测试上的性能可媲美当前最优方法,尤其在参考图像较少的稀疏场景中展现出显著更强的鲁棒性。

原文摘要 · Abstract (English)

Standard visual localization methods typically require offline pre-processing of scenes to obtain 3D structural information for better performance. This inevitably introduces additional computational and time costs, as well as the overhead of storing scene representations. Can we visually localize in a wild scene without any off-line preprocessing step? In this paper, we leverage the online inference capabilities of feed-forward 3D reconstruction networks to propose a novel map-free visual localization framework $L^3$. Specifically, by performing direct online 3D reconstruction on RGB images, followed by two-stage metric scale recovery and pose refinement based on 2D-3D correspondences, $L^3$ achieves high accuracy without the need to pre-build or store any offline scene representations. Extensive experiments demonstrate $L^3$ not only that the performance is comparable to state-of-the-art solutions on various benchmarks, but also that it exhibits significantly superior robustness in sparse scenes (fewer reference images per scene).

视觉定位3D重建免地图实时定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。