让机器人读懂导航标识,提升自主导航能力
Sign Language: Towards Sign Understanding for Robot Autonomy
- 用视觉语言模型解析标识中的位置与方向信息
- 在医院、商场等场景测试,准确率超基准方法
- 适合研究机器人环境理解与人机交互的学者
导航标识是人类路径规划和场景理解的重要辅助,但机器人对其利用不足。我们认为,这些标识直接编码了动作、空间区域及关系等关键信息,能显著提升机器人导航与场景理解能力。由于场景与标识复杂多变,开放世界下的标识理解仍具挑战,但近期视觉语言模型(VLMs)的发展使其成为可能。为此,我们提出导航标识理解任务,旨在从标识中解析出位置与关联方向。我们构建了该任务的基准,设计合适评估指标,并整理了一个包含多样化公共空间(如医院、商场、交通枢纽)中不同复杂度标识的测试集。同时提供基于VLM的基线方法,验证其在该任务上的潜力。代码与数据集已开源。
原文摘要 · Abstract (English)
Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understanding, by directly encoding privileged information on actions, spatial regions, and relations. Interpreting signs in open-world settings remains a challenge owing to the complexity of scenes and signs, but recent advances in vision-language models (VLMs) make this feasible. To advance progress in this area, we introduce the task of navigational sign understanding which parses locations and associated directions from signs. We offer a benchmark for this task, proposing appropriate evaluation metrics and curating a test set capturing signs with varying complexity and design across diverse public spaces, from hospitals to shopping malls to transport hubs. We also provide a baseline approach using VLMs, and demonstrate their promise on navigational sign understanding. Code and dataset are available on Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。