arXiv:2512.23870cs.LG2025-12

用流模型提升强化学习策略表达力,实现更优的最优动作分布学习。

Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR

  • 用流模型参数化策略,增强表达能力与鲁棒性。
  • 提出在线重要性采样流匹配算法,仅需用户指定采样分布即可更新策略。
  • 在LQR问题上验证可学习最优动作分布,理论分析揭示采样分布对效率的影响。

软演员-批评家(SAC)是最大熵强化学习中的主流算法。实践中,为提高效率常使用简单策略类近似能量基策略,牺牲了表达力和鲁棒性。本文提出一种SAC变体,采用流模型参数化策略,利用其强大表达能力。算法中通过瞬时变量变换技术评估流模型策略,并采用本文提出的在线流匹配变体进行策略更新。该在线变体称为重要性采样流匹配(ISFM),可在仅使用用户指定采样分布的情况下完成策略更新,无需目标分布信息。我们对ISFM进行了理论分析,揭示了不同采样分布对学习效率的影响。最后,在最大熵线性二次调节器(LQR)问题上进行案例研究,证明所提算法能学习到最优动作分布。

原文摘要 · Abstract (English)

Soft actor-critic (SAC) is a popular algorithm for max-entropy reinforcement learning. In practice, the energy-based policies in SAC are often approximated using simple policy classes for efficiency, sacrificing the expressiveness and robustness. In this paper, we propose a variant of the SAC algorithm that parameterizes the policy with flow-based models, leveraging their rich expressiveness. In the algorithm, we evaluate the flow-based policy utilizing the instantaneous change-of-variable technique and update the policy with an online variant of flow matching developed in this paper. This online variant, termed importance sampling flow matching (ISFM), enables policy update with only samples from a user-specified sampling distribution rather than the unknown target distribution. We develop a theoretical analysis of ISFM, characterizing how different choices of sampling distributions affect the learning efficiency. Finally, we conduct a case study of our algorithm on the max-entropy linear quadratic regulator problems, demonstrating that the proposed algorithm learns the optimal action distribution.

强化学习流模型SACLQR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。