arXiv:2503.01178cs.LGcs.RO2025-03AAAI被引 5

提升可微环境下的模型训练精度与策略学习稳定性

Differentiable Information Enhanced Model-Based Reinforcement Learning

  • 用梯度惩罚的索博列夫训练提升动态模型预测精度
  • 通过截断学习窗口混合长度降低策略梯度方差
  • 适合需要高精度控制的机器人任务,如人形机器人运动

可微环境为基于梯度的方法提供了丰富的可微信息,促进了控制策略的学习。相比主流的无模型强化学习方法,基于模型的强化学习(MBRL)有潜力更有效地利用这些信息来还原底层物理动力学。然而,这带来了两个主要挑战:如何有效利用可微信息以构建预测更准确的动力学模型,以及如何增强策略训练的稳定性。本文提出一种可微信息增强的MBRL方法——MB-MIX,以应对这两个问题。首先,采用索博列夫(Sobolev)模型训练方式,对错误的模型梯度输出施加惩罚,从而提升预测精度,获得更精确、忠实于系统动力学的模型。其次,引入截断学习窗口的混合长度机制,降低策略梯度估计的方差,显著提升策略学习过程的稳定性。我们在多个涉及可控刚体机器人(如人形机器人运动控制和柔体物体操作)的复杂任务中进行了理论分析与实验验证,结果表明该方法优于以往的基于模型及无模型方法。

原文摘要 · Abstract (English)

Differentiable environments have heralded new possibilities for learning control policies by offering rich differentiable information that facilitates gradient-based methods. In comparison to prevailing model-free reinforcement learning approaches, model-based reinforcement learning (MBRL) methods exhibit the potential to effectively harness the power of differentiable information for recovering the underlying physical dynamics. However, this presents two primary challenges: effectively utilizing differentiable information to 1) construct models with more accurate dynamic prediction and 2) enhance the stability of policy training. In this paper, we propose a Differentiable Information Enhanced MBRL method, MB-MIX, to address both challenges. Firstly, we adopt a Sobolev model training approach that penalizes incorrect model gradient outputs, enhancing prediction accuracy and yielding more precise models that faithfully capture system dynamics. Secondly, we introduce mixing lengths of truncated learning windows to reduce the variance in policy gradient estimation, resulting in improved stability during policy learning. To validate the effectiveness of our approach in differentiable environments, we provide theoretical analysis and empirical results. Notably, our approach outperforms previous model-based and model-free methods, in multiple challenging tasks involving controllable rigid robots such as humanoid robots' motion control and deformable object manipulation.

强化学习可微建模机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。