提出新回放策略,让机械臂用更少尝试学会复杂动作。
Next-Future: Sample-Efficient Policy Learning for Robotic-Arm Tasks
- 通过奖励单步转移提升学习效率
- 8个任务中7个样本效率显著提升
- 适合对精度要求高的机器人控制场景
Hindsight Experience Replay (HER) 是机器人操作任务中多目标强化学习的主流方法,适用于二元奖励环境。尽管其能从失败轨迹中学习,但依赖启发式回放策略,缺乏理论依据。为此,本文提出新回放策略 Next-Future,聚焦于单步转移的奖励设计,显著提升了多目标马尔可夫决策过程的学习效率与准确性,尤其在严苛精度要求下表现优异。实验在八个复杂机器人操作任务上验证,使用十组随机种子训练,结果表明:七项任务样本效率显著提升,六项任务成功率更高。真实世界实验进一步验证了所学策略的可行性,展示了 Next-Future 在解决复杂机械臂任务中的潜力。
原文摘要 · Abstract (English)
Hindsight Experience Replay (HER) is widely regarded as the state-of-the-art algorithm for achieving sample-efficient multi-goal reinforcement learning (RL) in robotic manipulation tasks with binary rewards. HER facilitates learning from failed attempts by replaying trajectories with redefined goals. However, it relies on a heuristic-based replay method that lacks a principled framework. To address this limitation, we introduce a novel replay strategy, "Next-Future", which focuses on rewarding single-step transitions. This approach significantly enhances sample efficiency and accuracy in learning multi-goal Markov decision processes (MDPs), particularly under stringent accuracy requirements -- a critical aspect for performing complex and precise robotic-arm tasks. We demonstrate the efficacy of our method by highlighting how single-step learning enables improved value approximation within the multi-goal RL framework. The performance of the proposed replay strategy is evaluated across eight challenging robotic manipulation tasks, using ten random seeds for training. Our results indicate substantial improvements in sample efficiency for seven out of eight tasks and higher success rates in six tasks. Furthermore, real-world experiments validate the practical feasibility of the learned policies, demonstrating the potential of "Next-Future" in solving complex robotic-arm tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。