让机器人像人一样读路牌问路,高效导航复杂建筑。
Human-like Navigation in a World Built for Humans
- 用视觉语言模型理解环境符号,决策时模拟人类推理。
- 在真实与仿真环境中实现高效导航,减少盲目搜索。
- 适合需要自主导航的智能机器人场景。
当在未访问过的室内环境(如办公楼)中导航时,人类会通过阅读标识和向他人询问等方式高效抵达目的地,从而避免大面积搜寻。现有机器人导航系统缺乏此类行为能力,导致在大型环境中效率低下。本文提出 ReasonNav,一个模块化导航系统,通过引入视觉语言模型(VLM)的推理能力,整合人类式导航技能。我们基于导航地标设计了紧凑的输入输出抽象,使 VLM 能聚焦于语言理解与推理。在真实与仿真任务上的评估表明,该代理能成功运用高层次推理,在大型复杂建筑中实现高效导航。
原文摘要 · Abstract (English)
When navigating in a man-made environment they haven't visited before--like an office building--humans employ behaviors such as reading signs and asking others for directions. These behaviors help humans reach their destinations efficiently by reducing the need to search through large areas. Existing robot navigation systems lack the ability to execute such behaviors and are thus highly inefficient at navigating within large environments. We present ReasonNav, a modular navigation system which integrates these human-like navigation skills by leveraging the reasoning capabilities of a vision-language model (VLM). We design compact input and output abstractions based on navigation landmarks, allowing the VLM to focus on language understanding and reasoning. We evaluate ReasonNav on real and simulated navigation tasks and show that the agent successfully employs higher-order reasoning to navigate efficiently in large, complex buildings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。