用多任务强化学习让机器人模仿人类动作并适应新场景。
Generalizing from References using a Multi-Task Reference and Goal-Driven RL Framework
- 将参考动作作为行为先验,而非强制约束。
- 在复杂地形上实现自然动作迁移,且保持动作流畅性。
- 适合需要灵活动作生成的机器人研究者。
从人类运动中学习敏捷人形机器人行为是一种实现自然协调控制的有效途径,但现有方法存在固有权衡:参考跟踪策略在演示数据集外往往脆弱,而纯任务驱动的强化学习虽具适应性却牺牲了动作质量。本文提出一种统一的多任务强化学习框架,将参考运动视为行为塑造的先验而非部署时的约束。单一目标条件策略同时训练两个任务,共享观测与动作空间,但在初始化方式、指令空间和奖励结构上不同:(i) 参考引导的模仿任务,参考轨迹提供密集模仿奖励但不作为策略输入;(ii) 目标条件化的泛化任务,目标独立于参考采样,奖励仅反映任务成功。通过在共享框架内联合优化,策略既从密集参考监督中获得结构化的人类动作技能,又能适应新目标与初始状态。该方法无需对抗目标、显式轨迹跟踪、相位变量或依赖参考的推理。我们在一个要求多样运动能力(如跳跃与攀爬)的箱体公园地形上评估,结果表明所学控制器可超越参考分布进行迁移,同时保持动作自然性。最后,通过组合多个习得技能,展示了长时程行为生成的能力,体现了策略在复杂场景中的灵活性。
原文摘要 · Abstract (English)
Learning agile humanoid behaviors from human motion offers a powerful route to natural, coordinated control, but existing approaches face a persistent trade-off: reference-tracking policies are often brittle outside the demonstration dataset, while purely task-driven Reinforcement Learning (RL) can achieve adaptability at the cost of motion quality. We introduce a unified multi-task RL framework that bridges this gap by treating reference motion as a prior for behavioral shaping rather than a deployment-time constraint. A single goal-conditioned policy is trained jointly on two tasks that share the same observation and action spaces, but differ in their initialization schemes, command spaces, and reward structures: (i) a reference-guided imitation task in which reference trajectories define dense imitation rewards but are not provided as policy inputs, and (ii) a goal-conditioned generalization task in which goals are sampled independently of any reference and where rewards reflect only task success. By co-optimizing these objectives within a shared formulation, the policy acquires structured, human-like motor skills from dense reference supervision while learning to adapt these skills to novel goals and initial conditions. This is achieved without adversarial objectives, explicit trajectory tracking, phase variables, or reference-dependent inference. We evaluate the method on a challenging box-based parkour playground that demands diverse athletic behaviors (e.g., jumping and climbing), and show that the learned controller transfers beyond the reference distribution while preserving motion naturalness. Finally, we demonstrate long-horizon behavior generation by composing multiple learned skills, illustrating the flexibility of the learned polices in complex scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。