用大模型自动把语言指令转成机器人操作,省去人工设计奖励函数。
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models
- 用GPT-4从自然语言描述自动生成奖励函数
- 在仿真中训练机器人完成单臂与双臂操作任务
- 支持真实机器人部署,实现端到端自动化
大型语言模型(LLM)和视觉语言模型(VLM)的发展推动了机器人高阶语义运动规划。强化学习(RL)可通过交互与奖励信号自主优化复杂行为,但设计有效奖励函数仍具挑战,尤其在现实任务中稀疏奖励不足、密集奖励需繁琐设计。本文提出自治强化学习框架ARCHIE,利用GPT-4从自然语言任务描述中直接生成奖励函数,用于模拟环境中的RL训练,并形式化奖励生成过程以提升可行性。同时,GPT-4自动编码任务成功标准,实现从文本到可部署机器人技能的一次性全自动转换。通过在ABB YuMi协作机器人上开展的单臂与双臂操作任务大量仿真实验验证了方法的有效性与实用性,部分任务已在真实机器人平台上成功演示。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and Visual Language Models (VLMs) have significantly impacted robotics, enabling high-level semantic motion planning applications. Reinforcement Learning (RL), a complementary paradigm, enables agents to autonomously optimize complex behaviors through interaction and reward signals. However, designing effective reward functions for RL remains challenging, especially in real-world tasks where sparse rewards are insufficient and dense rewards require elaborate design. In this work, we propose Autonomous Reinforcement learning for Complex Human-Informed Environments (ARCHIE), an unsupervised pipeline leveraging GPT-4, a pre-trained LLM, to generate reward functions directly from natural language task descriptions. The rewards are used to train RL agents in simulated environments, where we formalize the reward generation process to enhance feasibility. Additionally, GPT-4 automates the coding of task success criteria, creating a fully automated, one-shot procedure for translating human-readable text into deployable robot skills. Our approach is validated through extensive simulated experiments on single-arm and bi-manual manipulation tasks using an ABB YuMi collaborative robot, highlighting its practicality and effectiveness. Tasks are demonstrated on the real robot setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。