arXiv:2603.09237cs.RO2026-03

让机器人多目标强化学习快270倍,几分钟内找到最优策略组合。

MO-Playground: Massively Parallelized Multi-Objective Reinforcement Learning for Robotics

  • 基于GPU的新型算法MORLAX,支持大规模并行仿真。
  • 在6个真实目标上实现机器人行走策略的帕累托最优,速度提升25-270倍。
  • 适合需要快速优化多任务机器人的研究者与工程师使用。

多目标强化学习(MORL)是学习在冲突目标间权衡的帕累托最优策略家族的强大工具。然而,与传统强化学习不同,现有MORL算法未能有效利用大规模并行化来同时模拟数千个环境,导致计算时间大幅增加,限制了其在复杂多目标机器人问题中的应用。为此,我们提出:1)MORLAX,一种新的原生GPU、快速的MORL算法;2)MO-Playground,一个可pip安装的GPU加速多目标环境平台。两者结合可在数分钟内逼近帕累托集,相比传统基于CPU的方法提速25-270倍,同时获得更优的帕累托前沿超体积。我们通过MO-Playground构建了自定义的BRUCE人形机器人环境,实现了在6个真实目标(如平滑性、效率、手臂摆动)下的帕累托最优行走策略学习。

原文摘要 · Abstract (English)

Multi-objective reinforcement learning (MORL) is a powerful tool to learn Pareto-optimal policy families across conflicting objectives. However, unlike traditional RL algorithms, existing MORL algorithms do not effectively leverage large-scale parallelization to concurrently simulate thousands of environments, resulting in vastly increased computation time. Ultimately, this has limited MORL's application towards complex multi-objective robotics problems. To address these challenges, we present 1) MORLAX, a new GPU-native, fast MORL algorithm, and 2) MO-Playground, a pip-installable playground of GPU-accelerated multi-objective environments. Together, MORLAX and MO-Playground approximate Pareto sets within minutes, offering 25-270x speed-ups compared to legacy CPU-based approaches whilst achieving superior Pareto front hypervolumes. We demonstrate the versatility of our approach by implementing a custom BRUCE humanoid robot environment using MO-Playground and learning Pareto-optimal locomotion policies across 6 realistic objectives for BRUCE, such as smoothness, efficiency and arm swinging.

强化学习机器人多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。