arXiv:2511.08942cs.ROcs.AI2025-11被引 3

让视觉语言模型当导航大脑,自动规划路线

Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning

  • 用思维链提示让模型分步推理,制定导航策略
  • 结合动作历史和俯视地图,避免原地打转,路径更直接
  • 适合想用大模型提升机器人自主导航的开发者

尽管视觉语言模型(VLMs)有望变革机器人导航,但现有方法常未能充分发挥其推理能力。为释放VLM在机器人领域的潜力,我们将其角色从被动观察者转变为导航过程中的主动策略制定者。框架将高层规划任务交由VLM完成,利用其上下文理解能力指导基于前哨的探索代理。通过三种技术实现智能引导:结构化思维链提示以激发逻辑性分步推理;动态引入代理近期动作历史,防止陷入循环;以及一项新能力,使VLM能同时解析俯视障碍图与第一视角图像,增强空间感知。在HM3D、Gibson和MP3D等挑战性基准测试中,该方法生成了异常直接且逻辑清晰的轨迹,显著优于现有方法,在导航效率上实现大幅提升,为构建更强大的具身智能体指明方向。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) are set to transform robotic navigation, existing methods often underutilize their reasoning capabilities. To unlock the full potential of VLMs in robotics, we shift their role from passive observers to active strategists in the navigation process. Our framework outsources high-level planning to a VLM, which leverages its contextual understanding to guide a frontier-based exploration agent. This intelligent guidance is achieved through a trio of techniques: structured chain-of-thought prompting that elicits logical, step-by-step reasoning; dynamic inclusion of the agent's recent action history to prevent getting stuck in loops; and a novel capability that enables the VLM to interpret top-down obstacle maps alongside first-person views, thereby enhancing spatial awareness. When tested on challenging benchmarks like HM3D, Gibson, and MP3D, this method produces exceptionally direct and logical trajectories, marking a substantial improvement in navigation efficiency over existing approaches and charting a path toward more capable embodied agents.

机器人导航视觉语言模型思维链具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。