arXiv:2603.01047cs.LGstat.ML2026-03中稿 · ICLR被引 2

提出新评估方法,让生成模型训练更稳定灵活。

Evaluating GFlowNet from partial episodes for stable and flexible policy-based training

  • 用部分轨迹的流量平衡评估策略差异
  • 在真实任务上提升训练可靠性和灵活性
  • 适合需要高效采样的强化学习研究者

生成流网络(GFlowNets)通过将生成过程建模为有向无环图中的轨迹,学习高效采样组合候选的策略。现有基于价值的训练流程要求学习策略与目标策略在部分轨迹上的流量保持平衡,从而隐式最小化策略差异。而基于策略的方法需交替估计策略差异并更新策略,但在有向无环图中可靠估计差异仍是难点。本文揭示流量平衡可作为策略评估的理论基础,提出一种基于部分轨迹的评估平衡目标来学习评估器。在合成与真实任务上的实验表明,该评估平衡不仅增强了基于策略训练的可靠性,还通过支持参数化反向策略和融合离线数据收集技术,显著提升了训练灵活性。

原文摘要 · Abstract (English)

Generative Flow Networks (GFlowNets) were developed to learn policies for efficiently sampling combinatorial candidates by interpreting their generative processes as trajectories in directed acyclic graphs. In the value-based training workflow, the objective is to enforce the balance over partial episodes between the flows of the learned policy and the estimated flows of the desired policy, implicitly encouraging policy divergence minimization. The policy-based strategy alternates between estimating the policy divergence and updating the policy, but reliable estimation of the divergence under directed acyclic graphs remains a major challenge. This work bridges the two perspectives by showing that flow balance also yields a principled policy evaluator that measures the divergence, and an evaluation balance objective over partial episodes is proposed for learning the evaluator. As demonstrated on both synthetic and real-world tasks, evaluation balance not only strengthens the reliability of policy-based training but also broadens its flexibility by seamlessly supporting parameterized backward policies and enabling the integration of offline data-collection techniques.

生成模型强化学习策略评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。