arXiv:2502.02316cs.LG2025-02ICML被引 59

用扩散模型提升强化学习探索能力,性能显著优于现有方法。

DIME:Diffusion-Based Maximum Entropy Reinforcement Learning

  • 提出基于扩散模型的极大熵强化学习框架,解决熵不可计算难题。
  • 在高维控制任务上超越其他扩散模型方法,接近顶尖非扩散方法性能。
  • 算法设计更简洁,更新数据比更低,适合资源受限场景使用。

极大熵强化学习(MaxEnt-RL)因其优越的探索特性已成为标准方法。传统策略常采用高斯分布参数化,严重限制表达能力。扩散模型策略更具表现力,但融入MaxEnt-RL面临主要挑战——其边际熵难以计算。为此,我们提出扩散基极大熵强化学习(DIME)。DIME利用扩散模型近似推断的最新进展,推导出最大熵目标的下界,并提出一种可证明收敛到最优扩散策略的策略迭代方案。该方法在保留MaxEnt-RL原则性探索优势的同时,显著优于其他基于扩散模型的方法,在挑战性的高维控制基准测试中表现突出。同时,其性能与最先进非扩散方法相当,且所需算法设计选择更少、更新/数据比更低,降低了计算复杂度。

原文摘要 · Abstract (English)

Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which significantly limits their representational capacity. Diffusion-based policies offer a more expressive alternative, yet integrating them into MaxEnt-RL poses challenges-primarily due to the intractability of computing their marginal entropy. To overcome this, we propose Diffusion-Based Maximum Entropy RL (DIME). \emph{DIME} leverages recent advances in approximate inference with diffusion models to derive a lower bound on the maximum entropy objective. Additionally, we propose a policy iteration scheme that provably converges to the optimal diffusion policy. Our method enables the use of expressive diffusion-based policies while retaining the principled exploration benefits of MaxEnt-RL, significantly outperforming other diffusion-based methods on challenging high-dimensional control benchmarks. It is also competitive with state-of-the-art non-diffusion based RL methods while requiring fewer algorithmic design choices and smaller update-to-data ratios, reducing computational complexity.

强化学习扩散模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。