arXiv:2512.12793cs.RO2025-12

用带标签的脚印图实现机器人全局定位,无需精确几何信息。

VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps

  • 利用视觉语言模型匹配地图中的地标名称与图像
  • 在模拟和真实零售环境中小幅优于传统扫描方法
  • 适合无详细地图但有可读标记的复杂场景

本文提出一种新型全局定位方法VLG-Loc,使用仅包含环境内显著视觉地标名称和区域的人类可读标签脚印图。尽管人类能借助此类地图自主定位,但将该能力迁移至机器人系统仍极具挑战,因缺乏几何与外观细节难以建立观测地标与地图中地标的对应关系。VLG-Loc通过视觉语言模型(VLM)在机器人的多方向图像中搜索地图中标注的地标,并在蒙特卡洛定位框架中评估各姿态假设的似然性。实验在模拟与真实零售环境中验证,相比现有基于扫描的方法,在环境变化下表现出更强鲁棒性。通过概率融合视觉与扫描定位,进一步提升性能。

原文摘要 · Abstract (English)

This paper presents Vision-Language Global Localization (VLG-Loc), a novel global localization method that uses human-readable labeled footprint maps containing only names and areas of distinctive visual landmarks in an environment. While humans naturally localize themselves using such maps, translating this capability to robotic systems remains highly challenging due to the difficulty of establishing correspondences between observed landmarks and those in the map without geometric and appearance details. To address this challenge, VLG-Loc leverages a vision-language model (VLM) to search the robot's multi-directional image observations for the landmarks noted in the map. The method then identifies robot poses within a Monte Carlo localization framework, where the found landmarks are used to evaluate the likelihood of each pose hypothesis. Experimental validation in simulated and real-world retail environments demonstrates superior robustness compared to existing scan-based methods, particularly under environmental changes. Further improvements are achieved through the probabilistic fusion of visual and scan-based localization.

视觉定位多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。