通过模拟中的特权动作,让机器人高效学会复杂长时序操作技能。
Learning Long-Horizon Robot Manipulation Skills via Privileged Action
- 在仿真中引入特权动作增强物体交互与探索效率
- 无需复杂奖励设计即可完成多阶段长时序任务
- 适合需要鲁棒长时序操作的机器人研究者
长时序、接触密集型任务在强化学习中难以训练,因高维状态空间探索低效且奖励稀疏。学习过程常陷入局部最优,复杂场景需任务特定奖励调优。本文提出一种结构化框架,结合课程学习与特权动作,在无需大量奖励工程或参考轨迹的情况下,使策略高效习得长时序技能。具体地,在仿真中使用放宽约束与虚拟力等特权动作,提升与物体的交互与探索能力。实验成功实现融合非抓握与抓取的复杂多阶段任务,从非可抓取姿态抬起物体。通过简洁奖励结构,展现跨多种环境的多样化且鲁棒行为收敛。真实世界实验进一步验证所学技能可迁移,表现出稳健复杂的性能。该方法在多项任务上优于现有最优方法,收敛至其他方法失败的解。
原文摘要 · Abstract (English)
Long-horizon contact-rich tasks are challenging to learn with reinforcement learning, due to ineffective exploration of high-dimensional state spaces with sparse rewards. The learning process often gets stuck in local optimum and demands task-specific reward fine-tuning for complex scenarios. In this work, we propose a structured framework that leverages privileged actions with curriculum learning, enabling the policy to efficiently acquire long-horizon skills without relying on extensive reward engineering or reference trajectories. Specifically, we use privileged actions in simulation with a general training procedure that would be infeasible to implement in real-world scenarios. These privileges include relaxed constraints and virtual forces that enhance interaction and exploration with objects. Our results successfully achieve complex multi-stage long-horizon tasks that naturally combine non-prehensile manipulation with grasping to lift objects from non-graspable poses. We demonstrate generality by maintaining a parsimonious reward structure and showing convergence to diverse and robust behaviors across various environments. Additionally, real-world experiments further confirm that the skills acquired using our approach are transferable to real-world environments, exhibiting robust and intricate performance. Our approach outperforms state-of-the-art methods in these tasks, converging to solutions where others fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。