arXiv:2504.17636cs.CV2025-04IJCV被引 10

对比无结构视觉定位方法,发现几何推理越强精度越高。

A Guide to Structureless Visual Localization

  • 用图像库替代3D模型,实现灵活场景更新
  • 基于经典姿态估计的方法精度显著优于端到端回归
  • 适合需频繁更新场景的应用,如AR/VR和自动驾驶

视觉定位算法用于估计查询图像在已知场景中的相机位姿,是自动驾驶和增强/混合现实系统的核心。当前主流方法依赖3D场景模型,通过2D-3D对应关系进行位姿估计,虽精度高但难以适应场景变化。无结构方法将场景表示为带已知位姿的图像数据库,可轻松增删图像实现灵活更新。本文首次系统综述并比较了无结构方法,实验表明:采用更强经典几何推理的方法(如绝对或半通用相对位姿估计)显著优于近年流行的位姿回归方法。与顶尖结构化方法相比,无结构方法虽精度略低,但灵活性优势明显,为未来研究提供新方向。

原文摘要 · Abstract (English)

Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars and augmented / mixed reality systems. State-of-the-art visual localization algorithms are structure-based, i.e., they store a 3D model of the scene and use 2D-3D correspondences between the query image and 3D points in the model for camera pose estimation. While such approaches are highly accurate, they are also rather inflexible when it comes to adjusting the underlying 3D model after changes in the scene. Structureless localization approaches represent the scene as a database of images with known poses and thus offer a much more flexible representation that can be easily updated by adding or removing images. Although there is a large amount of literature on structure-based approaches, there is significantly less work on structureless methods. Hence, this paper is dedicated to providing the, to the best of our knowledge, first comprehensive discussion and comparison of structureless methods. Extensive experiments show that approaches that use a higher degree of classical geometric reasoning generally achieve higher pose accuracy. In particular, approaches based on classical absolute or semi-generalized relative pose estimation outperform very recent methods based on pose regression by a wide margin. Compared with state-of-the-art structure-based approaches, the flexibility of structureless methods comes at the cost of (slightly) lower pose accuracy, indicating an interesting direction for future work.

视觉定位无结构方法位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。