arXiv:2505.01822cs.LGcs.AI2025-05NeurIPS被引 2

用解析方法解决离线强化学习中能量引导的生成难题。

Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

  • 基于高斯条件变换,推导出中间能量的闭式解。
  • 在温和假设下估计对数期望,实现能量引导的精准控制。
  • 在30多个离线强化学习任务中超越主流基线方法。

条件生成决策的扩散模型在强化学习中展现出强大竞争力。最新研究揭示了能量函数引导的扩散模型与约束强化学习问题之间的关联。主要挑战在于生成过程中由于对数期望形式导致的中间能量难以估计。为此,我们提出解析能量引导策略优化(AEPO)。首先,在扩散模型服从条件高斯变换的前提下,提供理论分析并给出中间引导的闭式解。接着,分析对数期望形式中的后验高斯分布,在温和假设下获得对数期望的目标估计。最后,训练一个中间能量神经网络以逼近该目标估计。我们在30多个离线强化学习任务上验证了方法的有效性。大量实验表明,本方法在D4RL离线强化学习基准上优于多个代表性基线。

原文摘要 · Abstract (English)

Conditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-guidance diffusion models and constrained RL problems. The main challenge lies in estimating the intermediate energy, which is intractable due to the log-expectation formulation during the generation process. To address this issue, we propose the Analytic Energy-guided Policy Optimization (AEPO). Specifically, we first provide a theoretical analysis and the closed-form solution of the intermediate guidance when the diffusion model obeys the conditional Gaussian transformation. Then, we analyze the posterior Gaussian distribution in the log-expectation formulation and obtain the target estimation of the log-expectation under mild assumptions. Finally, we train an intermediate energy neural network to approach the target estimation of log-expectation formulation. We apply our method in 30+ offline RL tasks to demonstrate the effectiveness of our method. Extensive experiments illustrate that our method surpasses numerous representative baselines in D4RL offline reinforcement learning benchmarks.

强化学习扩散模型离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。