arXiv:2505.14139cs.LGcs.AI2025-05被引 10

用能量引导流模型训练,让离线强化学习无需推理时调参。

FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning

  • 通过能量引导路径近似为高斯路径,学习条件速度场作为策略
  • 在离线强化学习中达到可比性能,训练时间与采样步数无关
  • 适合需要多模态动作分布的复杂控制任务

在扩散模型中,利用引导机制使采样趋向期望结果已被广泛研究,尤其在图像和轨迹生成中。然而,将引导机制融入训练过程仍相对较少被探索。本文提出能量引导流匹配,一种新方法,用于增强流模型训练,并消除推理阶段对引导的需求。我们通过将能量引导概率路径近似为高斯路径,学习对应的条件速度场作为流策略。该方法适用于目标分布由数据与能量函数共同定义的任务,如强化学习。近年来,基于扩散的策略因其表达能力强、能捕捉多模态动作分布而受到关注。通常这些策略通过加权目标函数或反向传播策略采样的动作梯度进行优化。作为替代方案,我们提出FlowQ——一种基于能量引导流匹配的离线强化学习算法。该方法在保持竞争性表现的同时,策略训练时间与流采样步数无关。

原文摘要 · Abstract (English)

The use of guidance to steer sampling toward desired outcomes has been widely explored within diffusion models, especially in applications such as image and trajectory generation. However, incorporating guidance during training remains relatively underexplored. In this work, we introduce energy-guided flow matching, a novel approach that enhances the training of flow models and eliminates the need for guidance at inference time. We learn a conditional velocity field corresponding to the flow policy by approximating an energy-guided probability path as a Gaussian path. Learning guided trajectories is appealing for tasks where the target distribution is defined by a combination of data and an energy function, as in reinforcement learning. Diffusion-based policies have recently attracted attention for their expressive power and ability to capture multi-modal action distributions. Typically, these policies are optimized using weighted objectives or by back-propagating gradients through actions sampled by the policy. As an alternative, we propose FlowQ, an offline reinforcement learning algorithm based on energy-guided flow matching. Our method achieves competitive performance while the policy training time is constant in the number of flow sampling steps.

强化学习扩散模型流匹配离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。