arXiv:2410.18065cs.ROcs.AI2024-10被引 15

融合规划与强化学习,让机器人更高效完成复杂长程操作。

SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation

  • 先用规划分解任务,再结合模仿与强化学习解决子问题。
  • 任务成功率提升35%~50%,所需人类示范减少6倍。
  • 适合需要长时间、高接触精度的机器人操作场景。

机器人学习在编程机械臂方面已被证明是通用且高效的方法。模仿学习仅需人类示范即可教学,但受限于示范质量;强化学习通过探索发现更优行为,但可优化空间过大,难以从零开始。对于两者而言,任务越长,学习难度越高。为此,我们提出SPIRE:首先利用任务与运动规划(TAMP)将任务分解为更小的学习子问题,再结合模仿学习与强化学习以最大化各自优势。我们设计了新的策略,使学习代理能在规划系统中有效训练。我们在一系列长时程、高接触频率的机器人操作任务上评估SPIRE,结果表明其在平均任务性能上比现有集成模仿学习、强化学习与规划的方法提升35%至50%,训练所需人类示范数量减少6倍,且任务执行效率接近翻倍。

原文摘要 · Abstract (English)

Robot learning has proven to be a general and effective technique for programming manipulators. Imitation learning is able to teach robots solely from human demonstrations but is bottlenecked by the capabilities of the demonstrations. Reinforcement learning uses exploration to discover better behaviors; however, the space of possible improvements can be too large to start from scratch. And for both techniques, the learning difficulty increases proportional to the length of the manipulation task. Accounting for this, we propose SPIRE, a system that first uses Task and Motion Planning (TAMP) to decompose tasks into smaller learning subproblems and second combines imitation and reinforcement learning to maximize their strengths. We develop novel strategies to train learning agents when deployed in the context of a planning system. We evaluate SPIRE on a suite of long-horizon and contact-rich robot manipulation problems. We find that SPIRE outperforms prior approaches that integrate imitation learning, reinforcement learning, and planning by 35% to 50% in average task performance, is 6 times more data efficient in the number of human demonstrations needed to train proficient agents, and learns to complete tasks nearly twice as efficiently. View https://sites.google.com/view/spire-corl-2024 for more details.

机器人操作强化学习模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。