用语义对齐提升OpenStreetMap中单目重定位的精度与速度
Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment
- 基于DINO-ViT提取视觉特征,缩小图像与地图间语义差异
- 分层搜索:先快速傅里叶匹配粗筛,再按不确定性精调局部
- 在单一数据集训练下,3°朝向召回率超越现有方法5°标准
单目重定位使机器人能根据视觉观测估计相机位姿。然而,许多现有方法依赖密集地图或大型参考图像数据库,面临可扩展性限制和隐私风险。OpenStreetMap(OSM)作为轻量级、隐私友好的地图,具备全局可扩展性的语义与几何信息。但自然图像与OSM之间存在跨模态差异,且基于全局地图的定位成本高昂。本文提出一种具有不确定性感知的分层搜索框架,结合语义对齐实现OSM中的定位。首先,利用以物体为中心的DINO-ViT标记,缩小地面视图观测与OSM向量间的语义差距;其次,将全局密集匹配分解为粗粒度的FFT相关性匹配与受不确定性控制的局部精化。大量实验表明,该方法显著提升了定位精度与速度。在仅使用一个数据集训练的情况下,其3°方向召回率甚至超过当前最优方法的5°召回率。
原文摘要 · Abstract (English)
Monocular re-localization enables robots to estimate camera poses from visual observations. However, many existing methods rely on dense maps or large reference image databases, which face scalability limitations and privacy risks. OpenStreetMap (OSM), as a lightweight privacy-preserving map, offers semantic and geometric information with global scalability. Nonetheless, OSM localization remains challenging due to cross-modal discrepancies between natural images and OSM, as well as the high cost of global map-based localization. In this paper, we propose an uncertainty-aware hierarchical search framework with semantic alignment for localization in OSM. First, object-centric DINO-ViT tokens are exploited to reduce the semantic gap between ground-view observations and OSM vectors. Second, global dense matching is decomposed into coarse FFT correlation and uncertainty-controlled local refinement. Extensive experiments demonstrate that our method significantly improves localization accuracy and speed. When trained on a single dataset, the 3$^\circ$ orientation recall of our method even outperforms the 5$^\circ$ recall of state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。