用优化生成合理动作,让机器人更稳地完成复杂操作。
Opt2Skill: Imitating Dynamically-feasible Whole-Body Trajectories for Versatile Humanoid Loco-Manipulation
- 结合动态规划与强化学习,生成稳定全身运动轨迹。
- 实测跟踪精度和任务成功率均优于人类示范与传统方法。
- 适合需要精准力控的机器人操作场景,如擦桌子。
类人机器人需执行多样化的移动操作任务,但其高维不稳定动力学和复杂的接触环境带来挑战。基于模型的最优控制虽能定义精确运动,却受限于高计算成本和对接触感知的依赖;而强化学习虽具强鲁棒性,却存在学习效率低、动作不自然及仿真到现实的差距问题。为此,我们提出 Opt2Skill,一种端到端框架,融合模型驱动的轨迹优化与强化学习,实现稳健的全身移动操作。该方法利用微分动态规划(DDP)为 Digit 类人机器人生成动态可行且接触一致的参考轨迹,并训练强化学习策略追踪这些最优轨迹。结果表明,Opt2Skill 在运动跟踪和任务成功率上均优于依赖人类示范与逆运动学参考的基线方法。此外,引入含力矩信息的轨迹可显著提升接触任务(如擦桌子)中的接触力跟踪性能。该方法已成功应用于真实世界场景。
原文摘要 · Abstract (English)
Humanoid robots are designed to perform diverse loco-manipulation tasks. However, they face challenges due to their high-dimensional and unstable dynamics, as well as the complex contact-rich nature of the tasks. Model-based optimal control methods offer flexibility to define precise motion but are limited by high computational complexity and accurate contact sensing. On the other hand, reinforcement learning (RL) handles high-dimensional spaces with strong robustness but suffers from inefficient learning, unnatural motion, and sim-to-real gaps. To address these challenges, we introduce Opt2Skill, an end-to-end pipeline that combines model-based trajectory optimization with RL to achieve robust whole-body loco-manipulation. Opt2Skill generates dynamic feasible and contact-consistent reference motions for the Digit humanoid robot using differential dynamic programming (DDP) and trains RL policies to track these optimal trajectories. Our results demonstrate that Opt2Skill outperforms baselines that rely on human demonstrations and inverse kinematics-based references, both in motion tracking and task success rates. Furthermore, we show that incorporating trajectories with torque information improves contact force tracking in contact-involved tasks, such as wiping a table. We have successfully transferred our approach to real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。