arXiv:2506.17462cs.ROcs.AI2025-06被引 1

用大模型驱动机器人自主规划导航,无需预先建图

General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting

  • 基于大模型构建可自定义任务流程的智能体框架
  • 在未测绘环境中实现鲁棒导航,准确率超越现有方法
  • 适合研究通用机器人系统与智能体架构的学者

开发适用于未知环境的通用导航策略仍是机器人领域的核心挑战。现有系统多依赖特定任务的神经网络和固定信息流,限制了泛化能力。大型视觉语言模型(LVLM)通过嵌入类人知识,为推理与规划提供了新思路,但以往的集成方案仍依赖预建地图、硬编码表示和僵化控制逻辑。本文提出面向智能体的机器人导航架构(ARNA),赋予基于LVLM的智能体一套来自现代机器人工具栈的感知、推理与导航工具库。运行时,智能体可自主定义并执行任务定制的工作流,迭代调用模块、融合多模态输入进行推理,并选择导航动作。该智能体范式使系统在未测绘环境中具备强健的导航与推理能力,为机器人系统设计提供了新视角。在Habitat Lab的HM-EQA基准上评估,ARNA优于现有EQA专用方法。在RxR及自定义任务上的定性结果进一步证明其在广泛导航挑战中的泛化能力。

原文摘要 · Abstract (English)

Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges.

机器人导航大模型智能体通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。