让机器人通过解读路标实现无地图导航。
SignScene: Visual Sign Grounding for Mapless Navigation
- 设计符号中心的时空语义表示,帮助模型理解路标与场景关系。
- 在9类环境中测试,路标定位准确率达88%,显著优于基线。
- 可在真实机器人上仅靠路标完成无地图导航,适合智能机器人应用。
导航标志使人类能在陌生环境中无需地图进行导航。本文研究机器人如何在开放世界中类似地利用路标实现无地图导航。核心挑战在于解读路标:真实路标种类繁多、语义抽象,需将其语义内容与局部3D场景进行关联。我们提出“符号接地”(sign grounding)问题,即把路标上的语义指令映射到对应场景元素和导航动作。近期视觉-语言模型(VLMs)具备所需的常识与推理能力,但对空间信息表达方式敏感。为此,我们提出SignScene——一种以符号为中心的时空语义表示,捕捉与导航相关的场景元素和符号信息,并以利于有效推理的形式呈现给VLMs。我们在包含114个查询的跨九类环境数据集上评估该方法,取得88%的接地准确率,显著优于基线。最终,我们验证其可在Spot机器人上实现真实世界的无地图导航,仅依赖路标信息。
原文摘要 · Abstract (English)
Navigational signs enable humans to navigate unfamiliar environments without maps. This work studies how robots can similarly exploit signs for mapless navigation in the open world. A central challenge lies in interpreting signs: real-world signs are diverse and complex, and their abstract semantic contents need to be grounded in the local 3D scene. We formalize this as sign grounding, the problem of mapping semantic instructions on signs to corresponding scene elements and navigational actions. Recent Vision-Language Models (VLMs) offer the semantic common-sense and reasoning capabilities required for this task, but are sensitive to how spatial information is represented. We propose SignScene, a sign-centric spatial-semantic representation that captures navigation-relevant scene elements and sign information, and presents them to VLMs in a form conducive to effective reasoning. We evaluate our grounding approach on a dataset of 114 queries collected across nine diverse environment types, achieving 88% grounding accuracy and significantly outperforming baselines. Finally, we demonstrate that it enables real-world mapless navigation on a Spot robot using only signs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。