直接动态重定向让机器人更精准模仿人类动作
Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos

- 单阶段直接优化动作轨迹,跳过中间几何转换
- 在物理模拟器中生成高保真、可执行的运动路径
- 适合需要精准动作模仿与快速强化学习训练的场景
从单目视频演示中进行模仿学习为教授人形机器人复杂技能提供了一种可扩展的方法。然而,将人类动作迁移到人形机器人需克服显著的形态差异。现有方法依赖几何重定向或间接动态重定向流程,这些中间运动学投影引入几何偏差,限制搜索空间并导致次优动态行为。本文提出直接动态重定向(DDR),一种新颖的单阶段框架,能直接从专家视频生成高保真、动力学可行的轨迹。通过在任务空间中建模问题,并在物理模拟器中使用基于采样的模型预测控制求解器,DDR原生优化复杂接触序列的同时减轻输入漂移。实验表明,摒弃几何偏差使DDR在演示追踪精度上超越现有最优基线。此外,提供此类物理可行参考可加速强化学习训练收敛,并提升敏捷与平衡行为的最终执行效果。源代码将公开。
原文摘要 · Abstract (English)
Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches. Standard approaches rely on Geometric Retargeting or Indirect Dynamic Retargeting pipelines. We identify that these intermediate kinematic projections introduce a geometric bias, restricting the search space and yielding suboptimal dynamic behaviors. In this paper, we propose Direct Dynamic Retargeting (DDR), a novel single-stage framework that generates high-fidelity, dynamically feasible trajectories directly from expert videos. By formulating the problem in the task space and leveraging a sampling-based Model Predictive Control solver within a physics simulator, DDR natively optimizes over complex contact sequences while mitigating input drift. Our experiments demonstrate that bypassing the geometric bias allows DDR to outperform state-of-the-art baselines in demonstration tracking accuracy. Furthermore, we establish that providing such physically viable references to RL agents accelerates training convergence and enhances the final execution of agile and balancing behaviors. Source code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。