让机器人持续学习高动态动作,同时不丢掉已学技能。
Extreme-RGMT: Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control

- 分两阶段训练:先学通用动作,再通过不对称机制强化难动作。
- 在多种高动态动作上完成率显著提升,尤其罕见动作表现更好。
- 适合研究通用人形机器人控制与持续学习的学者参考。
人类能逐步掌握高度动态的运动技能,同时保持日常动作的可靠性。而现有仿人机器人控制器在通用性与专业性之间存在权衡:通用动作追踪策略难以可靠执行罕见的高动态动作,而专项训练又会损害已有行为。我们提出Extreme-RGMT,一种两阶段持续学习框架,用于实现鲁棒的通用人形机器人控制。该方法首先从多源异构运动数据中学习一个通用动作追踪基础策略,随后采用非对称技能获取与能力巩固机制,在约束已掌握动作漂移的同时,重点强化困难的动态片段。为应对高动态动作稀缺、高失败率及有效样本不足的问题,Extreme-RGMT结合难度感知采样与优势优先轨迹重采样,突出关键段落。实验表明,Extreme-RGMT在通用全身动作追踪任务上达到当前最优性能,尤其在挑战性高动态动作的完成率上大幅提升。所获控制器可直接在固定参考与在线惯性动捕输入下执行多样未见过的高动态动作,推动通用全身动作追踪控制器向人类专家级高动态能力迈进。
原文摘要 · Abstract (English)
Humans can progressively acquire highly dynamic motor skills while preserving reliable everyday motor abilities. In contrast, existing humanoid controllers face a trade-off between generalist and specialist capabilities: generalist motion tracking policies struggle to reliably execute rare highly dynamic motions, whereas specialist training can degrade previously acquired behaviors. We introduce Extreme-RGMT, a two-stage continual learning framework for robust generalist humanoid control. The method first learns a generalist motion-tracking base policy from diverse multi-source motion data, then employs an asymmetric skill acquisition and capability consolidation mechanism to constrain policy drift on mastered motions while emphasizing difficult dynamic segments. To address the scarcity of highly dynamic motions, their high failure rates, and the resulting shortage of informative samples, Extreme-RGMT combines difficulty-aware sampling with advantage-prioritized trajectory resampling to emphasize critical segments. Experiments show that Extreme-RGMT achieves state-of-the-art generalist whole-body motion-tracking performance, including substantially improved completion of challenging highly dynamic motions. The resulting controller directly executes diverse unseen highly dynamic motions under fixed references and online inertial motion-capture inputs, advancing generalist whole-body motion-tracking controllers toward highly dynamic motor capabilities at the human-expert level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。