用多目标强化学习让机器人动作更灵活,一策多用省调参
AMOR: Adaptive Character Control through Multi-Objective Reinforcement Learning
- 训练单一策略,通过调节权重控制多种行为
- 可快速切换动作风格,支持高动态机器人运动
- 适合需要快速适应新任务的机器人控制场景
强化学习已显著提升物理驱动角色的运动控制能力,但传统方法依赖冲突奖励函数的加权和,需大量调参。由于强化学习计算成本高,此迭代过程耗时且繁琐。此外,为保证真实世界表现,权重选择需克服仿真到现实的差距。为此,我们提出一种多目标强化学习框架,训练一个基于权重条件的单一策略,覆盖奖励权衡的帕累托前沿。训练后可自由调整权重,大幅缩短迭代时间。实验表明该框架支持机器人完成高动态运动,并在层级设置中实现高层策略动态选择权重,使策略能高效适应新任务,编码多样行为。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has significantly advanced the control of physics-based and robotic characters that track kinematic reference motion. However, methods typically rely on a weighted sum of conflicting reward functions, requiring extensive tuning to achieve a desired behavior. Due to the computational cost of RL, this iterative process is a tedious, time-intensive task. Furthermore, for robotics applications, the weights need to be chosen such that the policy performs well in the real world, despite inevitable sim-to-real gaps. To address these challenges, we propose a multi-objective reinforcement learning framework that trains a single policy conditioned on a set of weights, spanning the Pareto front of reward trade-offs. Within this framework, weights can be selected and tuned after training, significantly speeding up iteration time. We demonstrate how this improved workflow can be used to perform highly dynamic motions with a robot character. Moreover, we explore how weight-conditioned policies can be leveraged in hierarchical settings, using a high-level policy to dynamically select weights according to the current task. We show that the multi-objective policy encodes a diverse spectrum of behaviors, facilitating efficient adaptation to novel tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。