arXiv:2503.04975cs.LG2025-03ICLR被引 68

提出无需辅助模型的能量引导流匹配方法,提升离线强化学习性能。

Energy-Weighted Flow Matching for Offline Reinforcement Learning

  • 直接学习能量引导的流匹配,避免额外训练中间模型。
  • 在离线强化学习任务中显著提升算法性能。
  • 首个不依赖辅助模型的精确能量引导流匹配方法,适合强化学习研究者。

本文研究生成建模中的能量引导机制,目标分布定义为 $q(/mathbf x) /propto p(/mathbf x)/exp(-β/mathcal E(/mathbf x))$,其中 $p(/mathbf x)$ 为数据分布,$/mathcal E(/mathbf x)$ 为能量函数。现有方法通常需在扩散过程中引入辅助模型学习中间引导信号。为此,我们提出能量加权流匹配(EFM),一种广义扩散过程,可直接学习能量引导的流而无需辅助模型。理论分析表明,该方法能准确捕捉引导流。进一步将此方法扩展至能量加权扩散模型,并应用于离线强化学习,提出 Q 加权迭代策略优化(QIPO)算法。实验表明,所提 QIPO 算法在离线 RL 任务中表现更优。值得注意的是,该算法是首个独立于辅助模型的能量引导扩散模型,也是文献中首个精确的能量引导流匹配模型。

原文摘要 · Abstract (English)

This paper investigates energy guidance in generative modeling, where the target distribution is defined as $q(\mathbf x) \propto p(\mathbf x)\exp(-β\mathcal E(\mathbf x))$, with $p(\mathbf x)$ being the data distribution and $\mathcal E(\mathcal x)$ as the energy function. To comply with energy guidance, existing methods often require auxiliary procedures to learn intermediate guidance during the diffusion process. To overcome this limitation, we explore energy-guided flow matching, a generalized form of the diffusion process. We introduce energy-weighted flow matching (EFM), a method that directly learns the energy-guided flow without the need for auxiliary models. Theoretical analysis shows that energy-weighted flow matching accurately captures the guided flow. Additionally, we extend this methodology to energy-weighted diffusion models and apply it to offline reinforcement learning (RL) by proposing the Q-weighted Iterative Policy Optimization (QIPO). Empirically, we demonstrate that the proposed QIPO algorithm improves performance in offline RL tasks. Notably, our algorithm is the first energy-guided diffusion model that operates independently of auxiliary models and the first exact energy-guided flow matching model in the literature.

强化学习生成模型扩散模型流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。