arXiv:2503.08084cs.ROcs.AI2025-03AAAI被引 13

让机器人通过理解指令和实时感知,自主完成复杂长序列操作。

Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile Manipulation

  • 用大模型将指令转为带环境约束的规划问题
  • 实测成功率超80%,7项操作技能均能完成
  • 适合需要自主决策的移动操作机器人场景

实现人形机器人在真实环境中基于具身感知与理解能力的长时序移动操作规划,长期面临挑战。随着大语言模型(LLMs)的兴起,基于LLM的规划方法迅速发展,但多数依赖人工提供的文本表征或大量提示工程,难以量化理解环境,例如判断物体可操作性。为此,我们提出指令增强型长时序规划系统(IALP),该框架利用LLM结合实时传感器反馈生成可行且最优的动作,在闭环交互中融入环境具身知识。不同于以往方法,本工作通过抽象推理与接地机制,将用户指令扩展为带有环境约束的PDDL问题。在多个真实世界长时序任务中,每项任务包含7种不同操作技能,实验结果表明,IALP系统平均成功率超过80%。所提方法可作为高层规划器,借助多模态传感器输入,赋予机器人在非结构化环境中强大的自主能力。

原文摘要 · Abstract (English)

Enabling humanoid robots to perform long-horizon mobile manipulation planning in real-world environments based on embodied perception and comprehension abilities has been a longstanding challenge. With the recent rise of large language models (LLMs), there has been a notable increase in the development of LLM-based planners. These approaches either utilize human-provided textual representations of the real world or heavily depend on prompt engineering to extract such representations, lacking the capability to quantitatively understand the environment, such as determining the feasibility of manipulating objects. To address these limitations, we present the Instruction-Augmented Long-Horizon Planning (IALP) system, a novel framework that employs LLMs to generate feasible and optimal actions based on real-time sensor feedback, including grounded knowledge of the environment, in a closed-loop interaction. Distinct from prior works, our approach augments user instructions into PDDL problems by leveraging both the abstract reasoning capabilities of LLMs and grounding mechanisms. By conducting various real-world long-horizon tasks, each consisting of seven distinct manipulatory skills, our results demonstrate that the IALP system can efficiently solve these tasks with an average success rate exceeding 80%. Our proposed method can operate as a high-level planner, equipping robots with substantial autonomy in unstructured environments through the utilization of multi-modal sensor inputs.

机器人规划大模型具身智能长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。