arXiv:2608.06021cs.RO2026-08

用视觉嵌入与3D模型结合,实现高精度自动驾驶定位。

Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models

论文配图:Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models
图 1 · 摘自论文原文
  • 将视觉定位与前馈3D几何模型融合,提升定位精度。
  • 在三个基准上显著优于现有基于外观的方法。
  • 适合需要抗视觉混淆的自动驾驶场景使用。

有效的视觉定位(VL)需要环境地图兼具紧凑性以支持高效扩展、对视觉变化具有鲁棒性以及具备度量精度。通过低维图像嵌入,视觉位置识别(VPR)能够满足前两项要求,但其度量精度较低,不如基于局部特征或神经表示的标准VL方法。该局限可通过将VPR与前馈神经3D几何(FF3D)模型产生的精确轨迹估计相结合来克服。本文提出一种拓扑-度量框架,通过受控图像集迭代结合概率VPR与FF3D度量位姿估计,实现序列化外观定位。该方法设计了自动离线映射工具,建模场景不同部分中位姿与外观的交互关系。该地图随后被在线粒子滤波器使用,结合里程计与对位置的信念进行FF3D推理,成功将神经度量估计融入概率外观定位中。我们在三个已知基准上广泛评估该框架,结果表明其显著优于现有外观定位方法。该方法的模块化设计使描述符提取器与FF3D模型可互换,进一步分析显示,序列信念可缓解严重感知伪影下的失败问题。

原文摘要 · Abstract (English)

Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirements, but its low metric accuracy makes it less suitable than standard VL approaches based on local features or neural representations. This limitation can be overcome by integrating VPR with the accurate local trajectory estimates produced by feed-forward neural 3D geometry (FF3D) models. In this paper, we address sequential appearance-based localization through a topometric framework that iteratively combines probabilistic VPR with FF3D metric pose estimation in controlled image sets. Our approach proposes an automatic offline mapping tool that models the topometric pose-appearance interaction in the different parts of the scene. This map is later employed by an online particle filter that estimates the pose from odometry and belief over places for FF3D inference, successfully incorporating neural metric estimation into probabilistic appearance-based localization. We extensively evaluate the framework on three known benchmarks, demonstrating substantial improvements over existing appearance-based methods. The modularity of our approach allows the descriptor extractor and FF3D model to remain interchangeable, and a focused analysis further shows that sequential belief can mitigate severe failures under perceptual aliasing.

自动驾驶视觉定位3D建模神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。