用手绘地图让机器人在不准确地图下也能导航
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
- 用视觉语言模型解析手绘地图,自动补全缺失信息
- 在仿真和真实场景中均实现90%以上导航成功率
- 适合家庭、办公室等复杂环境的通用导航需求
手绘地图能以自然高效的方式在人与机器人之间传递导航指令,但常存在比例失真、地标缺失等问题。本文提出一种新型手绘地图导航架构HAM-Nav,利用预训练视觉语言模型(VLMs)实现跨环境、多机器人形态下的导航,即使面对地图不准确仍有效。HAM-Nav采用选择性视觉关联提示方法进行拓扑地图定位与路径规划,并引入预测导航计划解析器推断缺失地标。在逼真的仿真环境中,使用轮式和腿式机器人进行了大量实验,验证了其在导航成功率和路径长度加权成功率方面的有效性。此外,在真实环境中的用户研究显示,手绘地图在实际导航中具有实用价值,表现优于非手绘地图方案。
原文摘要 · Abstract (English)
Hand-drawn maps can be used to convey navigation instructions between humans and robots in a natural and efficient manner. However, these maps can often contain inaccuracies such as scale distortions and missing landmarks which present challenges for mobile robot navigation. This paper introduces a novel Hand-drawn Map Navigation (HAM-Nav) architecture that leverages pre-trained vision language models (VLMs) for robot navigation across diverse environments, hand-drawing styles, and robot embodiments, even in the presence of map inaccuracies. HAM-Nav integrates a unique Selective Visual Association Prompting approach for topological map-based position estimation and navigation planning as well as a Predictive Navigation Plan Parser to infer missing landmarks. Extensive experiments were conducted in photorealistic simulated environments, using both wheeled and legged robots, demonstrating the effectiveness of HAM-Nav in terms of navigation success rates and Success weighted by Path Length. Furthermore, a user study in real-world environments highlighted the practical utility of hand-drawn maps for robot navigation as well as successful navigation outcomes compared against a non-hand-drawn map approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。