arXiv:2410.15474cs.LG2024-10ICLR被引 11

提升生成模型的逆向策略,加速复杂环境下的模式发现。

Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization

  • 直接最大化熵正则化MDP中的价值函数优化逆向策略
  • 在多个基准上实现更快收敛和更好模式覆盖
  • 适合需要高效探索的生成建模与强化学习场景

生成流网络(GFlowNets)是一类生成模型,通过学习按给定奖励函数比例采样对象。其核心思想是使用两个随机策略:前向策略逐步构建组合对象,逆向策略则逐步分解它们。近期研究发现,GFlowNet训练与熵正则化强化学习(RL)问题存在紧密联系,但该联系仅适用于固定逆向策略的情况,这可能成为显著限制。为解决此问题,我们提出一种简单的逆向策略优化算法,通过在中间奖励的熵正则化马尔可夫决策过程(MDP)中直接最大化价值函数实现。我们在多种基准上对所提方法进行了全面实验,结合了强化学习与GFlowNet算法,结果表明该方法在复杂环境中实现了更快的收敛速度和更优的模式发现能力。

原文摘要 · Abstract (English)

Generative Flow Networks (GFlowNets) are a family of generative models that learn to sample objects with probabilities proportional to a given reward function. The key concept behind GFlowNets is the use of two stochastic policies: a forward policy, which incrementally constructs compositional objects, and a backward policy, which sequentially deconstructs them. Recent results show a close relationship between GFlowNet training and entropy-regularized reinforcement learning (RL) problems with a particular reward design. However, this connection applies only in the setting of a fixed backward policy, which might be a significant limitation. As a remedy to this problem, we introduce a simple backward policy optimization algorithm that involves direct maximization of the value function in an entropy-regularized Markov Decision Process (MDP) over intermediate rewards. We provide an extensive experimental evaluation of the proposed approach across various benchmarks in combination with both RL and GFlowNet algorithms and demonstrate its faster convergence and mode discovery in complex environments.

生成模型强化学习策略优化流网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。