arXiv:2505.23458cs.LG2025-05被引 55

用扩散模型引导提升策略,无需价值函数就能让离线强化学习更好

Diffusion Guidance Is a Controllable Policy Improvement Operator

  • 将扩散模型的引导机制直接转化为策略改进算子
  • 引导权重越大,离线任务表现越好,且性能持续提升
  • 无需学习价值函数,可免费提升行为克隆等方法的最优性

强化学习的核心在于超越数据中的表现,但扩展此类系统一直非常困难。相比之下,生成建模技术表现出极强的可扩展性且易于训练。本文通过建立策略改进与扩散模型引导之间的直接关系,提出新框架CFGRL:它以监督学习的简单方式训练,却能进一步优化数据中的策略。在离线强化学习任务中,我们观察到明确趋势——引导权重增加,性能随之提升。尤为重要的是,CFGRL无需显式学习价值函数,从而可将简单的监督方法(如目标条件行为克隆)推广至更优策略,实现性能的‘免费’提升。

原文摘要 · Abstract (English)

At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.

强化学习扩散模型离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。