提出新方法提升人形机器人动作迁移质量,减少错误并提高控制鲁棒性。
Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- 设计通用动作迁移算法GMR,优化人类动作向机器人适配的准确性。
- 实验显示现有方法因动作误差导致策略鲁棒性下降,动态动作更受影响。
- GMR在追踪精度和动作忠实度上接近闭源数据,适合无调参场景使用。
人形机器人运动跟踪策略是构建遥操作流程与分层控制器的核心,但面临人类与机器人本体之间的表征差距问题。当前方法通过将人类动作数据迁移到人形机器人本体,并训练强化学习(RL)策略模仿参考轨迹来解决。然而,迁移过程中引入的伪影(如脚滑、自穿透、物理不可行动作)常被保留在参考轨迹中,需由RL策略自行修正。尽管已有工作展示出运动追踪能力,但通常依赖大量奖励工程与领域随机化才能成功。本文系统评估了迁移质量对策略性能的影响,且在抑制过度奖励调参的前提下进行测试。针对现有迁移方法的问题,我们提出一种新方法——通用动作迁移(GMR)。在与两个开源迁移器(PHC、ProtoMotions)及一个高质量闭源数据集(Unitree)对比的基础上,采用BeyondMimic进行策略训练,剥离奖励调参影响。在LAFAN1数据集的多样化子集上的实验表明,尽管多数动作可被追踪,但迁移数据中的伪影显著降低策略鲁棒性,尤其在动态或长序列动作中。GMR在追踪表现与源动作忠实度上均优于现有开源方法,其感知保真度与策略成功率接近闭源基准。
原文摘要 · Abstract (English)
Humanoid motion tracking policies are central to building teleoperation pipelines and hierarchical controllers, yet they face a fundamental challenge: the embodiment gap between humans and humanoid robots. Current approaches address this gap by retargeting human motion data to humanoid embodiments and then training reinforcement learning (RL) policies to imitate these reference trajectories. However, artifacts introduced during retargeting, such as foot sliding, self-penetration, and physically infeasible motion are often left in the reference trajectories for the RL policy to correct. While prior work has demonstrated motion tracking abilities, they often require extensive reward engineering and domain randomization to succeed. In this paper, we systematically evaluate how retargeting quality affects policy performance when excessive reward tuning is suppressed. To address issues that we identify with existing retargeting methods, we propose a new retargeting method, General Motion Retargeting (GMR). We evaluate GMR alongside two open-source retargeters, PHC and ProtoMotions, as well as with a high-quality closed-source dataset from Unitree. Using BeyondMimic for policy training, we isolate retargeting effects without reward tuning. Our experiments on a diverse subset of the LAFAN1 dataset reveal that while most motions can be tracked, artifacts in retargeted data significantly reduce policy robustness, particularly for dynamic or long sequences. GMR consistently outperforms existing open-source methods in both tracking performance and faithfulness to the source motion, achieving perceptual fidelity and policy success rates close to the closed-source baseline. Website: https://jaraujo98.github.io/retargeting_matters. Code: https://github.com/YanjieZe/GMR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。