人机协同强化学习让机器人快速掌握高精度操作技能
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
- 结合人类示范与修正,用视觉强化学习训练机器人
- 1-2.5小时达成近满分成功率,速度比之前快1.8倍
- 适合工业自动化与复杂操作研究者参考
强化学习(RL)在实现自主复杂机器人操作技能方面潜力巨大,但在真实环境中应用仍具挑战。本文提出一种基于视觉的人机协同强化学习系统,在动态操作、精密装配和双臂协作等多种精细操作任务中表现出色。该方法融合人类示范与修正、高效强化学习算法及系统级设计,仅需1至2.5小时训练即实现接近完美的成功率达,并显著缩短周期时间。实验表明,相比模仿学习基线和以往强化学习方法,平均成功率提升2倍,执行速度加快1.8倍。通过大量实验分析,揭示了该方法如何学习出具备鲁棒性与适应性的反应式与预测式控制策略。结果表明,强化学习可在实际环境中高效学习多样化的视觉驱动操作策略,为工业应用与科研进步提供新范式。视频与代码见项目网站 https://hil-serl.github.io/。
原文摘要 · Abstract (English)
Reinforcement learning (RL) holds great promise for enabling autonomous acquisition of complex robotic manipulation skills, but realizing this potential in real-world settings has been challenging. We present a human-in-the-loop vision-based RL system that demonstrates impressive performance on a diverse set of dexterous manipulation tasks, including dynamic manipulation, precision assembly, and dual-arm coordination. Our approach integrates demonstrations and human corrections, efficient RL algorithms, and other system-level design choices to learn policies that achieve near-perfect success rates and fast cycle times within just 1 to 2.5 hours of training. We show that our method significantly outperforms imitation learning baselines and prior RL approaches, with an average 2x improvement in success rate and 1.8x faster execution. Through extensive experiments and analysis, we provide insights into the effectiveness of our approach, demonstrating how it learns robust, adaptive policies for both reactive and predictive control strategies. Our results suggest that RL can indeed learn a wide range of complex vision-based manipulation policies directly in the real world within practical training times. We hope this work will inspire a new generation of learned robotic manipulation techniques, benefiting both industrial applications and research advancements. Videos and code are available at our project website https://hil-serl.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。