arXiv:2606.15793cs.LGcs.AI2026-06

将近端策略优化用于生成流网络,提升采样效率与收敛速度

Proximal Policy Optimization for Amortized Discrete Sampling

论文配图:Proximal Policy Optimization for Amortized Discrete Sampling
图 1 · 摘自论文原文
  • 基于生成流网络框架,推导出近端策略优化算法
  • 在合成能量与分子图生成任务中,收敛更快、数据利用率更高
  • 首次成功应用PPO于GFlowNet,适合结构化离散采样研究者

本文研究在生成流网络(GFlowNet)框架下,通过策略梯度算法训练随机策略以从结构化离散概率分布中采样。基于GFlowNets与熵正则强化学习之间的理论联系,我们推导了标准策略梯度算法在训练GFlowNets中的等价形式,并实验性探索了基线训练与优势估计等方法论问题。最重要的是,本工作首次推导并成功应用近端策略优化(PPO)于GFlowNets,在从合成能量到分子图生成的多个基准测试中,相比传统GFlowNet训练目标展现出更优的收敛速度与数据效率。

原文摘要 · Abstract (English)

This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distributions under the Generative Flow Network (GFlowNet) framework. Building on extensive theoretical connections between GFlowNets and entropy-regularized reinforcement learning, we derive equivalents of standard policy gradient algorithms for training GFlowNets, as well as experimentally explore their various methodological aspects, including baseline training and advantage estimation. Most importantly, our work is the first to derive and successfully apply proximal policy optimization to GFlowNets, showing its improved convergence speed and data efficiency compared to standard GFlowNet training objectives on benchmarks ranging from synthetic energies to molecular graph generation.

生成流网络策略优化离散采样PPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。