用4小时数据训练出能稳定控制人形机器人做高动态动作的通用运动追踪策略。
EGM: Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control
- 按动作难度动态采样,提升训练效率
- 分体处理上下肢特征,增强对不同动作的适应性
- 仅用4.08小时数据即在49.25小时测试中表现优异
从人类动作中学习通用运动追踪策略在人形机器人全身控制中具有巨大潜力。传统方法在数据利用和训练效率上表现不佳,且难以追踪高动态动作。为此,我们提出EGM框架,包含四项核心设计:首先引入基于动作分组的跨动作课程自适应采样策略,根据各动作分组的追踪误差动态调整采样概率,高效平衡不同难度与持续时间动作的训练;所采数据由提出的复合解耦专家混合(CDMoE)架构处理,通过分别对上肢和下肢分组专家,并将正交专家与共享专家解耦,分别处理专用特征与通用特征,显著提升对不同分布动作的追踪能力;关键洞察是:训练通用运动追踪策略时,数据质量与多样性至关重要。在此基础上,我们设计三阶段课程式训练流程,逐步提升策略对扰动的鲁棒性。尽管仅使用4.08小时数据训练,EGM在49.25小时测试动作中展现出稳健泛化能力,优于基线模型,在常规与高动态任务中均表现更优。
原文摘要 · Abstract (English)
Learning a general motion tracking policy from human motions shows great potential for versatile humanoid whole-body control. Conventional approaches are not only inefficient in data utilization and training processes but also exhibit limited performance when tracking highly dynamic motions. To address these challenges, we propose EGM, a framework that enables efficient learning of a general motion tracking policy. EGM integrates four core designs. Firstly, we introduce a Bin-based Cross-motion Curriculum Adaptive Sampling strategy to dynamically orchestrate the sampling probabilities based on tracking error of each motion bin, eficiently balancing the training process across motions with varying dificulty and durations. The sampled data is then processed by our proposed Composite Decoupled Mixture-of-Experts (CDMoE) architecture, which efficiently enhances the ability to track motions from different distributions by grouping experts separately for upper and lower body and decoupling orthogonal experts from shared experts to separately handle dedicated features and general features. Central to our approach is a key insight we identified: for training a general motion tracking policy, data quality and diversity are paramount. Building on these designs, we develop a three-stage curriculum training flow to progressively enhance the policy's robustness against disturbances. Despite training on only 4.08 hours of data, EGM generalized robustly across 49.25 hours of test motions, outperforming baselines on both routine and highly dynamic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。