用单次示范训练机器人完成复杂多接触操作,还能自动修复失误。
Guided Reinforcement Learning for Robust Multi-Contact Loco-Manipulation
- 用任务无关的强化学习框架,仅凭一次示教就能训练策略。
- 在真实机器人上实现高成功率,能自动生成重新抓取等补救动作。
- 适合需要灵活应对干扰的工业机器人任务,如推门、搬重物。
强化学习通常需要为每项任务精心设计马尔可夫决策过程。本文提出一种系统性方法,用于多接触运动与操作任务的行为合成与控制,例如穿越弹簧门和操控重型洗碗机。通过定义一个任务无关的马尔可夫决策过程,利用基于模型的轨迹优化器生成的单次示范进行策略训练。方法引入自适应相位动力学形式,使策略能稳健跟踪示范,同时适应动态不确定性与外部扰动。与先前运动模仿强化学习方法对比,所学策略在所有任务中均达到更高成功率。策略还学会了示范中未包含的恢复动作,如执行中的重新抓取或处理滑移。最终,该策略成功部署于真实机器人,验证了方法的实际可行性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) often necessitates a meticulous Markov Decision Process (MDP) design tailored to each task. This work aims to address this challenge by proposing a systematic approach to behavior synthesis and control for multi-contact loco-manipulation tasks, such as navigating spring-loaded doors and manipulating heavy dishwashers. We define a task-independent MDP to train RL policies using only a single demonstration per task generated from a model-based trajectory optimizer. Our approach incorporates an adaptive phase dynamics formulation to robustly track the demonstrations while accommodating dynamic uncertainties and external disturbances. We compare our method against prior motion imitation RL works and show that the learned policies achieve higher success rates across all considered tasks. These policies learn recovery maneuvers that are not present in the demonstration, such as re-grasping objects during execution or dealing with slippages. Finally, we successfully transfer the policies to a real robot, demonstrating the practical viability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。