让机器人在真实世界学会灵活抓取,成功率近92%。
Dexterous Grasping with Real-World Robotic Reinforcement Learning
- 先用少量专家示范预训练,再在真实环境强化学习微调。
- 真实抓取成功率接近92%,平均耗时减少23%。
- 适合需要真实场景部署的灵巧操作研究者。
真实世界中的灵巧抓取是机器人学习的核心挑战。能够根据物体特性在任意场景中灵活抓取不同几何与属性的物体,对通用机器人至关重要。然而,现有研究多在仿真环境中进行,因现实与仿真之间的域差距,难以直接应用于真实场景,限制了其泛化性与实用性。本文提出 DexGraspRL,一种直接在真实环境中训练机器人获取灵巧抓取技能的强化学习框架。该框架包含两个阶段:(i) 使用少量专家示范通过模仿学习(IL)进行预训练;(ii) 在真实场景中通过强化学习(RL)进行微调。为缓解演示数据与真实环境间分布偏移导致的灾难性遗忘,设计了正则化项,平衡强化学习的探索与预训练策略的保留。实验表明,DexGraspRL成功完成多种灵巧抓取任务,平均成功率达92%。通过强化学习微调,新策略比模仿学习策略平均周期时间减少23%。
原文摘要 · Abstract (English)
Dexterous grasping in the real world presents a fundamental and significant challenge for robot learning. The ability to employ affordance-aware poses to grasp objects with diverse geometries and properties in arbitrary scenarios is essential for general-purpose robots. However, existing research predominantly addresses dexterous grasping problems within simulators, which encounter difficulties when applied in real-world environments due to the domain gap between reality and simulation. This limitation hinders their generalizability and practicality in real-world applications. In this paper, we present DexGraspRL, a reinforcement learning (RL) framework that directly trains robots in real-world environments to acquire dexterous grasping skills. Specifically, DexGraspRL consists of two stages: (i) a pretraining stage that pretrains the policy using imitation learning (IL) with a limited set of expert demonstrations; (ii) a fine-tuning stage that refines the policy through direct RL in real-world scenarios. To mitigate the catastrophic forgetting phenomenon arising from the distribution shift between demonstrations and real-world environments, we design a regularization term that balances the exploitation of RL with the preservation of the pretrained policy. Our experiments with real-world tasks demonstrate that DexGraspRL successfully accomplishes diverse dexterous grasping tasks, achieving an average success rate of nearly 92%. Furthermore, by fine-tuning with RL, our method uncovers novel policies, surpassing the IL policy with a 23% reduction in average cycle time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。