让机器人听懂自然语言指令,自主完成移动与操作任务。
Grounding Language Models in Autonomous Loco-manipulation Tasks
- 用强化学习+全身优化生成动作,存入动作库
- 通过大模型构建分层任务图,实现从指令到执行的规划
- 在真实机器人上验证,能应对开放场景的自由文本指令
具有行为自主性的类人机器人被视为日常生活的理想协作伙伴和具身智能的有力体现。相比固定基座机械臂,类人机器人虽拥有更大的操作空间,但控制与规划难度显著提升。尽管通用类人机器人进展迅速,多数研究仍集中于行走能力,对全身协调与任务规划关注较少,限制了其在开放语境下执行长时序移动-操作任务的潜力。本文提出一种新框架,基于不同场景的任务学习、选择并规划行为。结合强化学习(RL)与全身优化生成机器人运动,并存入动作库;进一步利用大语言模型(LLM)的规划与推理能力,构建包含一系列运动基元的分层任务图,实现底层执行与高层规划之间的桥梁。仿真与真实世界实验(使用CENTAURO机器人)表明,基于语言模型的规划器可高效适应新移动-操作任务,在非结构化场景中仅凭自由文本命令即展现高度自主性。
原文摘要 · Abstract (English)
Humanoid robots with behavioral autonomy have consistently been regarded as ideal collaborators in our daily lives and promising representations of embodied intelligence. Compared to fixed-based robotic arms, humanoid robots offer a larger operational space while significantly increasing the difficulty of control and planning. Despite the rapid progress towards general-purpose humanoid robots, most studies remain focused on locomotion ability with few investigations into whole-body coordination and tasks planning, thus limiting the potential to demonstrate long-horizon tasks involving both mobility and manipulation under open-ended verbal instructions. In this work, we propose a novel framework that learns, selects, and plans behaviors based on tasks in different scenarios. We combine reinforcement learning (RL) with whole-body optimization to generate robot motions and store them into a motion library. We further leverage the planning and reasoning features of the large language model (LLM), constructing a hierarchical task graph that comprises a series of motion primitives to bridge lower-level execution with higher-level planning. Experiments in simulation and real-world using the CENTAURO robot show that the language model based planner can efficiently adapt to new loco-manipulation tasks, demonstrating high autonomy from free-text commands in unstructured scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。