arXiv:2602.12829cs.LGcs.AI2026-02被引 5

用动能正则化实现无需密度估计的最大熵强化学习

FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching

  • 将策略优化建模为广义薛定谔桥问题,以动能作为熵正则化代理
  • 在高维连续控制任务上性能优于或相当主流基线方法
  • 适合追求高效、免密度估计强化学习的算法研究者

迭代生成策略(如扩散模型和流匹配)在连续控制中具有更强表达能力,但使最大熵强化学习复杂化,因其动作对数密度不可直接获取。为此,我们提出场最小能量演员-评论家(FLAC),一种无需似然性的框架,通过惩罚速度场的动能来调节策略随机性。核心思想是将策略优化视为相对于高熵参考过程(如均匀分布)的广义薛定谔桥问题。在此视角下,最大熵原则自然出现:在优化回报的同时尽量贴近高熵参考,无需显式动作密度。动能作为物理上合理的参考偏差代理,最小化路径空间能量可约束最终动作分布的偏离。基于此,我们推导出能量正则化策略迭代方案,并设计了一种通过拉格朗日对偶机制自动调节动能的实用离策略算法。实验表明,FLAC在高维基准测试中表现优于或相当主流基线,同时避免了显式密度估计。

原文摘要 · Abstract (English)

Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning because their action log-densities are not directly accessible. To address this, we propose Field Least-Energy Actor-Critic (FLAC), a likelihood-free framework that regulates policy stochasticity by penalizing the kinetic energy of the velocity field. Our key insight is to formulate policy optimization as a Generalized Schrödinger Bridge (GSB) problem relative to a high-entropy reference process (e.g., uniform). Under this view, the maximum-entropy principle emerges naturally as staying close to a high-entropy reference while optimizing return, without requiring explicit action densities. In this framework, kinetic energy serves as a physically grounded proxy for divergence from the reference: minimizing path-space energy bounds the deviation of the induced terminal action distribution. Building on this view, we derive an energy-regularized policy iteration scheme and a practical off-policy algorithm that automatically tunes the kinetic energy via a Lagrangian dual mechanism. Empirically, FLAC achieves superior or comparable performance on high-dimensional benchmarks relative to strong baselines, while avoiding explicit density estimation.

强化学习扩散模型最大熵无密度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。