用仿真预训练提升真实机器人强化学习的样本效率。
SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training
- 先在数字孪生环境预训练视觉运动策略,再用于真实世界
- 样本效率显著提升,多阶段操作任务接近完美成功率
- 适合做复杂抓取与长时序操控的机器人研究者参考
自主学习灵巧、长时序的机器人技能是具身智能的长期目标。尽管近期机器人强化学习在真实世界的视觉运动控制任务中表现出色,但其应用仍面临样本效率低、探索缓慢和高度依赖人工干预等问题。相比之下,模拟器提供安全高效的探索与数据收集环境,而通过真实到仿真技术可缓解视觉仿真-真实差距。基于此,我们提出SimLauncher框架,结合真实世界强化学习与真实→仿真→真实方法的优势以克服上述挑战。具体而言,首先在数字孪生仿真环境中预训练一个视觉运动策略,该策略通过两种方式助力真实世界学习:(1) 利用大量仿真示范及由预训练策略生成的真实世界示范来初始化目标值;(2) 引入预训练策略的动作建议以提升探索质量。我们在多阶段、高接触性、灵巧手操作任务上进行了全面实验。相较于以往真实世界强化学习方法,SimLauncher显著提升了样本效率,并实现了接近完美的成功率。本工作为利用大规模仿真预训练促进真实机器人强化学习提供了概念验证,有望激发后续研究。
原文摘要 · Abstract (English)
Autonomous learning of dexterous, long-horizon robotic skills has been a longstanding pursuit of embodied AI. Recent advances in robotic reinforcement learning (RL) have demonstrated remarkable performance and robustness in real-world visuomotor control tasks. However, applying RL in the real world faces challenges such as low sample efficiency, slow exploration, and significant reliance on human intervention. In contrast, simulators offer a safe and efficient environment for extensive exploration and data collection, while the visual sim-to-real gap, often a limiting factor, can be mitigated using real-to-sim techniques. Building on these, we propose SimLauncher, a novel framework that combines the strengths of real-world RL and real-to-sim-to-real approaches to overcome these challenges. Specifically, we first pre-train a visuomotor policy in the digital twin simulation environment, which then benefits real-world RL in two ways: (1) bootstrapping target values using extensive simulated demonstrations and real-world demonstrations derived from pre-trained policy rollouts, and (2) Incorporating action proposals from the pre-trained policy for better exploration. We conduct comprehensive experiments across multi-stage, contact-rich, and dexterous hand manipulation tasks. Compared to prior real-world RL approaches, SimLauncher significantly improves sample efficiency and achieves near-perfect success rates. We hope this work serves as a proof of concept and inspires further research on leveraging large-scale simulation pre-training to benefit real-world robotic RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。