arXiv:2607.24409cs.CVcs.RO2026-07

用高精度街景影像实现厘米级视觉定位,验证了其媲美专业测绘的潜力。

Accuracy potential of visual localization exploiting high-end street-level imagery

论文配图:Accuracy potential of visual localization exploiting high-end street-level imagery
图 1 · 摘自论文原文
  • 直接以精确定位的高清街景图作场景表示,结合候选筛选与实时重建估姿。
  • 定位中位误差达1-5厘米(平移)和0.05-0.1度(旋转),最优时可达1厘米和0.03度。
  • 适合研究高精度定位、自动驾驶、数字孪生及消费级设备自动地理配准的团队。

在自动驾驶、测绘、机器人和增强/混合现实等应用中,对参考坐标系下的精确可靠位姿需求日益增长。视觉定位可作为GNSS的补充,但其精度潜力尚未系统评估,主要受限于缺乏公开的大规模室外数据集,且地面真值位姿需达到亚厘米级。本文填补了这一空白:提出一种可扩展的视觉定位流程,直接使用精确地理参考的高分辨率街景影像作为场景表示,结合先验引导的候选选择、在线结构光重建与基于PnP的位姿估计。同时发布FHNW Muttenz数据集,覆盖连续10公里道路网络,由两次相隔约1.5年的移动测绘采集,包含四种不同相机拍摄的五个代表性场景的高分辨率参考与查询图像序列。所有图像均精确配准,提供亚厘米级的6自由度地面真值位姿。实验表明,定位中位误差为1-5厘米(平移)和0.05-0.1度(旋转),在理想条件下最低可达1厘米和0.03度。结果证明视觉定位可媲美专业测绘级GNSS,为消费级设备实现三维地理空间数据采集与全自动地理配准铺平道路。数据集已公开:https://fhnw-muttenz-vl-dataset.github.io/

原文摘要 · Abstract (English)

Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicability and accuracy are often limited. Yet, the accuracy potential of visual localization has not been systematically investigated against survey-grade demands. This is mainly due to the lack of publicly available, large-scale outdoor datasets with ground-truth poses in the sub-centimeter range. In this work, we address both gaps. We introduce a scalable visual localization pipeline that employs precisely georeferenced, high-resolution street-level imagery directly as the scene representation. It combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation. We further present the FHNW Muttenz dataset, a real-world dataset covering a contiguous 10 km street network mapped in two mobile mapping campaigns approximately 1.5 years apart. It consists of high-resolution reference imagery and query sequences acquired by four different cameras across five representative scenes. All images are precisely co-registered, yielding 6-DoF ground-truth poses in the sub-centimeter range. Using this dataset, we evaluate the accuracy potential of visual localization. Our experiments demonstrate median pose accuracies in the range of 1-5 cm for translation and 0.05-0.1° for rotation, reaching as low as 1 cm and 0.03° under favorable conditions. These results show that visual localization can complement survey-grade GNSS positioning, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches. The dataset is publicly available at: https://fhnw-muttenz-vl-dataset.github.io/.

视觉定位高精度街景影像6D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。