arXiv:2608.03967stat.MLcs.LG2026-08

用信息几何优化生成流网络的前向策略,提升采样效率与探索能力。

Information-Geometric Forward Policy Training in GFlowNets

论文配图:Information-Geometric Forward Policy Training in GFlowNets
图 1 · 摘自论文原文
  • 基于轨迹采样器的几何结构,提出自然梯度更新方法。
  • 揭示了每步条件二阶矩对轨迹费舍尔信息的分解机制。
  • 适用于具有局部结构或可分解目标的生成任务,如图模型场景。

生成流网络(GFlowNets)是一种灵活的框架,用于对离散及混合离散-连续对象进行近似推断,仅需通过奖励函数定义非归一化目标密度。本文从诱导轨迹采样器的信息几何出发,形式化了前向策略的训练过程。将前向策略视为轨迹采样器,发现其一阶内在几何由轨迹族的费舍尔-罗信息度量决定,当费舍尔信息可计算或准确近似时,对应的自然梯度提供最优局部更新。我们推导出轨迹费舍尔信息的精确分解:每步条件二阶矩之和。该分解阐明了时间得分相互作用消失的条件,以及共享参数下密集耦合仍存在的情形。由此形成三种计算范式:费舍尔信息可解析计算、蒙特卡洛估计足够有效的场景,以及目标局域性或因子分解结构可被利用的场合。后者中,精确边缘化、分割器方法与信念传播等图模型工具可作为自然梯度更新的合理近似。新框架将目标结构转化为优化几何,为结构感知的前向策略训练提供了可行路径。实验验证了黎曼与欧氏优化在收敛速度与探索行为上的差异。

原文摘要 · Abstract (English)

Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical local update whenever the corresponding Fisher information is computable or accurately approximable. We derive an exact decomposition of the trajectory Fisher into per-step conditional second moments, which clarifies when temporal score interactions vanish and when dense couplings remain under shared parameterisation. This leads to three computational regimes: settings with tractable exact Fisher information, settings where Monte Carlo estimators of the expected Fisher are sufficient, and structure-exploitable settings in which target locality or factorisation yields accurate approximations of the Fisher expectation. In the latter case, graphical-model tools such as exact marginalisation, separator methods, and belief propagation provide principled surrogates for natural-gradient updates. The resulting framework turns target structure into optimisation geometry and yields a tractable route to structure-aware forward-policy training in GFlowNets. We illustrate the framework empirically through examples comparing convergence and exploration behaviour under Riemannian and Euclidean optimisation.

生成模型信息几何强化学习图模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。