arXiv:2412.12089cs.LGcs.AI2024-12ICLR被引 24

让强化学习在软体物理仿真中更稳定,提升复杂任务控制能力

Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation

  • 提出SAPO算法,利用可微分仿真的一阶梯度优化策略
  • 在包含刚体与柔体交互的任务上,性能优于现有基线方法
  • 配套开发Rewarped平台,支持多物理场并行仿真

基于GPU的并行仿真使从业者能在消费级显卡上收集大量数据,并使用深度强化学习训练复杂控制策略。然而,强化学习在机器人领域的成功主要局限于能用快速刚体动力学模拟的任务。软体仿真技术相比之下慢几个数量级,导致强化学习因样本复杂性要求受限。本文提出一种新型强化学习算法和仿真平台,以实现涉及刚体与可变形体任务的强化学习规模化。我们引入软体解析策略优化(SAPO),一种最大熵的一阶模型基演员-评论家算法,利用可微分仿真的一阶解析梯度,训练随机演员以最大化期望回报和熵。同时,我们开发了Rewarped,一个支持多种材料仿真的并行可微分多物理场仿真平台。我们在Rewarped中重新实现多个具有挑战性的操作与运动任务,结果表明,SAPO在涉及刚体、关节和可变形体交互的多种任务中均优于基线方法。更多信息见https://rewarped.github.io/。

原文摘要 · Abstract (English)

Recent advances in GPU-based parallel simulation have enabled practitioners to collect large amounts of data and train complex control policies using deep reinforcement learning (RL), on commodity GPUs. However, such successes for RL in robotics have been limited to tasks sufficiently simulated by fast rigid-body dynamics. Simulation techniques for soft bodies are comparatively several orders of magnitude slower, thereby limiting the use of RL due to sample complexity requirements. To address this challenge, this paper presents both a novel RL algorithm and a simulation platform to enable scaling RL on tasks involving rigid bodies and deformables. We introduce Soft Analytic Policy Optimization (SAPO), a maximum entropy first-order model-based actor-critic RL algorithm, which uses first-order analytic gradients from differentiable simulation to train a stochastic actor to maximize expected return and entropy. Alongside our approach, we develop Rewarped, a parallel differentiable multiphysics simulation platform that supports simulating various materials beyond rigid bodies. We re-implement challenging manipulation and locomotion tasks in Rewarped, and show that SAPO outperforms baselines over a range of tasks that involve interaction between rigid bodies, articulations, and deformables. Additional details at https://rewarped.github.io/.

强化学习可微分仿真软体物理多物理场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。