用仿真训练机器人完成复杂抓取任务,真实成功率近80%。
SLIM: Sim-to-Real Legged Instructive Manipulation via Long-Horizon Visuomotor Learning
- 分层策略:高层视觉指令决策+底层四足行走控制
- 教师-学生强化学习,实现长时序任务分解与视觉导航
- 仅用普通硬件,在室内外多种光照下成功部署
我们提出一种低成本的四足移动操作机器人系统,通过纯仿真强化学习训练,完成长时序真实世界任务。系统采用三层设计:1)高层策略负责根据任务指令进行视觉-移动-操作,低层策略控制四足行走;2)教师-学生训练框架,教师利用任务分解与目标物体信息攻克长时序任务,学生则在教师引导下通过强化学习训练视觉-移动-操作能力;3)一系列减少仿真到现实差距的技术。相比以往依赖高成本设备的工作,本系统使用Unitree Go1四足机器人、WidowX-250S机械臂和单个腕装RGB相机,在仿真中完全训练,单一策略即可自主完成搜索、移动、抓取、运输、放置等长时序任务,真实世界成功率接近80%。该性能接近专家人类遥控操作水平,且机器人效率更高,速度约为遥控操作的1.5倍。我们还进行了大量消融实验,验证了高效强化学习训练与有效仿真到现实迁移的关键技术,并在多种室内外场景及不同光照条件下实现了稳定部署。
原文摘要 · Abstract (English)
We present a low-cost legged mobile manipulation system that solves long-horizon real-world tasks, trained by reinforcement learning purely in simulation. This system is made possible by 1) a hierarchical design of a high-level policy for visual-mobile manipulation following task instructions, and a low-level quadruped locomotion policy, 2) a teacher and student training pipeline for the high level, which trains a teacher to tackle long-horizon tasks using privileged task decomposition and target object information, and further trains a student for visual-mobile manipulation via RL guided by the teacher's behavior, and 3) a suite of techniques for minimizing the sim-to-real gap. In contrast to many previous works that use high-end equipments, our system demonstrates effective performance with more accessible hardware -- specifically, a Unitree Go1 quadruped, a WidowX-250S arm, and a single wrist-mounted RGB camera -- despite the increased challenges of sim-to-real transfer. Trained fully in simulation, a single policy autonomously solves long-horizon tasks involving search, move to, grasp, transport, and drop into, achieving nearly 80% real-world success. This performance is comparable to that of expert human teleoperation on the same tasks while the robot is more efficient, operating at about 1.5x the speed of the teleoperation. Finally, we perform extensive ablations on key techniques for efficient RL training and effective sim-to-real transfer, and demonstrate effective deployment across diverse indoor and outdoor scenes under various lighting conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。