用NMPC生成演示数据,提升机器人抓取柔体的强化学习效率
Robot Deformable Object Manipulation via NMPC-generated Demonstrations in Deep Reinforcement Learning
- 结合NMPC自动生成高质量演示数据,降低人工标注成本
- 相比基线算法,平均奖励提升2.01倍,标准差降至45%
- 在真实场景中完成折叠与展平任务,成功率最高达100%
本研究基于示范增强的强化学习(RL)方法,探索机器人对柔体物体的操控问题。为提升学习效率,提出HGCR-DDPG算法:采用高维模糊抓取点选择策略,改进Rainbow-DDPG中的数据驱动学习,并引入序列化策略学习机制。相比基线算法(Rainbow-DDPG),HGCR-DDPG实现全局平均奖励提升2.01倍,标准差降至45%。为降低示范收集的人工成本,提出基于非线性模型预测控制(NMPC)的低成本示范生成方法。仿真结果表明,通过NMPC生成的示范可有效训练HGCR-DDPG,性能接近人工示范。物理实验中,对布料执行对角折叠、中心轴折叠和展平三项任务,成功率达83.3%、80%和100%,验证了方法的有效性。相比当前大型模型方法,该算法轻量化、资源消耗少,具备任务定制性和高效适应能力。
原文摘要 · Abstract (English)
In this work, we conducted research on deformable object manipulation by robots based on demonstration-enhanced reinforcement learning (RL). To improve the learning efficiency of RL, we enhanced the utilization of demonstration data from multiple aspects and proposed the HGCR-DDPG algorithm. It uses a novel high-dimensional fuzzy approach for grasping-point selection, a refined behavior-cloning method to enhance data-driven learning in Rainbow-DDPG, and a sequential policy-learning strategy. Compared to the baseline algorithm (Rainbow-DDPG), our proposed HGCR-DDPG achieved 2.01 times the global average reward and reduced the global average standard deviation to 45% of that of the baseline algorithm. To reduce the human labor cost of demonstration collection, we proposed a low-cost demonstration collection method based on Nonlinear Model Predictive Control (NMPC). Simulation experiment results show that demonstrations collected through NMPC can be used to train HGCR-DDPG, achieving comparable results to those obtained with human demonstrations. To validate the feasibility of our proposed methods in real-world environments, we conducted physical experiments involving deformable object manipulation. We manipulated fabric to perform three tasks: diagonal folding, central axis folding, and flattening. The experimental results demonstrate that our proposed method achieved success rates of 83.3%, 80%, and 100% for these three tasks, respectively, validating the effectiveness of our approach. Compared to current large-model approaches for robot manipulation, the proposed algorithm is lightweight, requires fewer computational resources, and offers task-specific customization and efficient adaptability for specific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。