用无监督预训练让机器人高效安全地学会复杂操作
SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training
- 用低精度仿真器无监督学习通用动作潜空间
- 真实世界1小时内学会高自由度全身操作任务
- 无需示范或人工设计先验,适合复杂机器人控制
构建具备能力的家用与工业机器人需掌握多自由度系统(如双臂移动机械臂)的控制。强化学习虽有潜力自主获取控制策略,但扩展至高自由度系统仍具挑战。直接在真实世界进行强化学习需兼顾探索安全性和样本效率,实践中难以实现;而模拟到现实的强化学习常因现实差距导致性能不稳定。本文提出SLAC,通过低精度模拟器预训练一个与任务无关的潜在动作空间,使真实世界强化学习成为可能。SLAC采用定制的无监督技能发现方法,促进时间抽象、解耦与安全性,从而提升下游学习效率。一旦学习到潜在动作空间,便作为新型离策略强化学习算法的动作接口,通过真实世界交互自主学习下游任务。我们在一系列双臂移动操作任务上评估了SLAC,其表现达到当前最优。值得注意的是,SLAC仅需不到一小时的真实世界交互即可学会富含接触的全身任务,且不依赖任何示范或手工设计的行为先验。更多信息及机器人视频见 robo-rl.github.io
原文摘要 · Abstract (English)
Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomously acquiring robot control policies, scaling it to high-DoF embodiments remains challenging. Direct RL in the real world demands both safe exploration and high sample efficiency, which are difficult to achieve in practice. Sim-to-real RL, on the other hand, is often brittle due to the reality gap. This paper introduces SLAC, a method that renders real-world RL feasible for complex embodiments by leveraging a low-fidelity simulator to pretrain a task-agnostic latent action space. SLAC trains this latent action space via a customized unsupervised skill discovery method designed to promote temporal abstraction, disentanglement, and safety, thereby facilitating efficient downstream learning. Once a latent action space is learned, SLAC uses it as the action interface for a novel off-policy RL algorithm to autonomously learn downstream tasks through real-world interactions. We evaluate SLAC against existing methods on a suite of bimanual mobile manipulation tasks, where it achieves state-of-the-art performance. Notably, SLAC learns contact-rich whole-body tasks in under an hour of real-world interactions, without relying on any demonstrations or hand-crafted behavior priors. More information and robot videos at robo-rl.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。