arXiv:2410.05364cs.LGcs.AI2024-10被引 45

用扩散模型做多步动作和动态建模,实现高效在线控制。

Diffusion Model Predictive Control

  • 用扩散模型学习多步动作与系统动态
  • 在D4RL上性能优于传统模型基规划方法
  • 可实时优化新奖励函数,适应新动态

我们提出扩散模型预测控制(D-MPC),一种新型的模型预测控制方法。该方法利用扩散模型同时学习多步动作提案和多步系统动态模型,并将其结合用于在线控制。在流行的D4RL基准测试中,D-MPC的表现显著优于现有的基于模型的离线规划方法(如MBOP),并达到与最先进的基于模型及无模型强化学习方法相当的水平。此外,我们展示了D-MPC在运行时优化新奖励函数、适应新动态的能力,并凸显其相较于现有基于扩散模型的规划基线的优势。

原文摘要 · Abstract (English)

We propose Diffusion Model Predictive Control (D-MPC), a novel MPC approach that learns a multi-step action proposal and a multi-step dynamics model, both using diffusion models, and combines them for use in online MPC. On the popular D4RL benchmark, we show performance that is significantly better than existing model-based offline planning methods using MPC (e.g. MBOP) and competitive with state-of-the-art (SOTA) model-based and model-free reinforcement learning methods. We additionally illustrate D-MPC's ability to optimize novel reward functions at run time and adapt to novel dynamics, and highlight its advantages compared to existing diffusion-based planning baselines.

扩散模型强化学习控制在线规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。