用大模型驱动机器人认知,实现家务任务的规划与自适应执行。
From Language to Action: Can LLM-Based Agents Be Used for Embodied Robot Cognition?
- 以大模型为核心,结合记忆模块构建机器人认知架构。
- 在模拟家庭环境中完成物品摆放与交换任务,表现出现有适应能力。
- 适合研究具身智能与大模型融合的开发者参考。
为在日常环境中灵活行动,机器人需具备多种认知能力以进行规划推理并执行恢复。虽然大语言模型(LLMs)展现出推理与语言理解等涌现能力,但将其用于控制具身机器人仍需可靠地将高层语言指令映射到低层感知与控制功能。本文探讨了大模型作为认知机器人架构中规划与执行推理核心的可行性。提出一种认知架构,其中基于代理的大模型负责核心规划与推理,工作记忆与情景记忆组件支持经验学习与适应。该架构用于控制一个移动操作机械臂,在模拟家庭环境中通过一系列高层工具(感知、推理、导航、抓取、放置)实现环境交互。在两个家务任务(物品放置与物品交换)上评估系统,考察其推理、规划与记忆利用能力。结果表明,大模型驱动的代理可完成结构化任务,并表现出涌现的适应性与记忆引导的规划能力,但也暴露出显著缺陷:对任务成功状态产生幻觉,以及拒绝承认并完成序列任务,导致指令遵循不佳。这些发现揭示了将大模型作为自主机器人具身认知控制器的潜力与挑战。
原文摘要 · Abstract (English)
In order to flexibly act in an everyday environment, a robotic agent needs a variety of cognitive capabilities that enable it to reason about plans and perform execution recovery. Large language models (LLMs) have been shown to demonstrate emergent cognitive aspects, such as reasoning and language understanding; however, the ability to control embodied robotic agents requires reliably bridging high-level language to low-level functionalities for perception and control. In this paper, we investigate the extent to which an LLM can serve as a core component for planning and execution reasoning in a cognitive robot architecture. For this purpose, we propose a cognitive architecture in which an agentic LLM serves as the core component for planning and reasoning, while components for working and episodic memories support learning from experience and adaptation. An instance of the architecture is then used to control a mobile manipulator in a simulated household environment, where environment interaction is done through a set of high-level tools for perception, reasoning, navigation, grasping, and placement, all of which are made available to the LLM-based agent. We evaluate our proposed system on two household tasks (object placement and object swapping), which evaluate the agent's reasoning, planning, and memory utilisation. The results demonstrate that the LLM-driven agent can complete structured tasks and exhibits emergent adaptation and memory-guided planning, but also reveal significant limitations, such as hallucinations about the task success and poor instruction following by refusing to acknowledge and complete sequential tasks. These findings highlight both the potential and challenges of employing LLMs as embodied cognitive controllers for autonomous robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。