arXiv:2602.01749cs.AIcs.LG2026-02被引 1

通过马尔可夫链视角改进生成流网络的探索与利用平衡。

Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives

  • 引入可调参数α,动态控制生成过程中的探索与利用策略。
  • 在分子生成等任务中,模式发现数量最多提升10倍。
  • 理论严谨,适合研究生成模型训练机制的学者使用。

生成流网络(GFlowNet)的目标隐含地固定了前向与后向策略的均衡混合,可能限制训练过程中的探索-利用权衡。通过进一步探讨GFlowNet与马尔可夫链的关联,我们建立了GFlowNet目标与马尔可夫链可逆性之间的等价关系,揭示了此类约束的根源,并提供了一个将马尔可夫链性质适配到GFlowNets的框架。基于这些理论成果,我们提出α-GFNs,通过可调参数α实现混合方式的泛化。该方法可直接调控探索-利用动态以增强模式发现能力,同时保证收敛至唯一流。在集合、位序列和分子生成等多个基准测试中,α-GFN目标均优于以往的GFlowNet目标,模式发现数量最多提升10倍。

原文摘要 · Abstract (English)

Generative Flow Network (GFlowNet) objectives implicitly fix an equal mixing of forward and backward policies, potentially constraining the exploration-exploitation trade-off during training. By further exploring the link between GFlowNets and Markov chains, we establish an equivalence between GFlowNet objectives and Markov chain reversibility, thereby revealing the origin of such constraints, and provide a framework for adapting Markov chain properties to GFlowNets. Building on these theoretical findings, we propose $α$-GFNs, which generalize the mixing via a tunable parameter $α$. This generalization enables direct control over exploration-exploitation dynamics to enhance mode discovery capabilities, while ensuring convergence to unique flows. Across various benchmarks, including Set, Bit Sequence, and Molecule Generation, $α$-GFN objectives consistently outperform previous GFlowNet objectives, achieving up to a $10 \times$ increase in the number of discovered modes.

生成模型马尔可夫链探索利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。