arXiv:2604.12909cs.RO2026-04

提出树状结构框架,让机器人持续学习新技能不遗忘。

Tree Learning: A Multi-Skill Continual Learning Framework for Humanoid Robots

  • 用根-分支参数继承机制复用运动先验,防止遗忘。
  • 支持周期与非周期动作,多模态适配提升训练效率。
  • 适合需实时交互的仿人机器人多技能部署场景。

随着强化学习在仿人机器人中从单任务向多技能演进,如何高效扩展新技能并避免灾难性遗忘成为具身智能的关键挑战。现有方法或依赖混合专家模型(MoE)复杂的拓扑调整,或需训练超大规模模型,难以轻量部署。为此,我们提出树状学习(Tree Learning)框架,采用根-分支层次化参数继承机制,通过参数复用为分支技能提供运动先验,从根本上防止遗忘。设计多模态前馈适应机制,结合相位调制与插值,支持周期性与非周期性运动。同时提出任务级奖励塑造策略,加速技能收敛。基于Unity的仿真实验表明,相比同时多任务训练,树状学习在多种代表性行走技能上获得更高奖励,且保持100%技能保留率,实现无缝多技能切换与实时交互控制。进一步在两个不同仿真任务中验证:一个类超级马里奥互动场景和经典中式园林自主导航环境,均展示出优异性能与泛化能力。

原文摘要 · Abstract (English)

As reinforcement learning for humanoid robots evolves from single-task to multi-skill paradigms, efficiently expanding new skills while avoiding catastrophic forgetting has become a key challenge in embodied intelligence. Existing approaches either rely on complex topology adjustments in Mixture-of-Experts (MoE) models or require training extremely large-scale models, making lightweight deployment difficult. To address this, we propose Tree Learning, a multi-skill continual learning framework for humanoid robots. The framework adopts a root-branch hierarchical parameter inheritance mechanism, providing motion priors for branch skills through parameter reuse to fundamentally prevent catastrophic forgetting. A multi-modal feedforward adaptation mechanism combining phase modulation and interpolation is designed to support both periodic and aperiodic motions. A task-level reward shaping strategy is also proposed to accelerate skill convergence. Unity-based simulation experiments show that, in contrast to simultaneous multi-task training, Tree Learning achieves higher rewards across various representative locomotion skills while maintaining a 100% skill retention rate, enabling seamless multi-skill switching and real-time interactive control. We further validate the performance and generalization capability of Tree Learning on two distinct Unity-simulated tasks: a Super Mario-inspired interactive scenario and autonomous navigation in a classical Chinese garden environment.

持续学习仿人机器人运动控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。