arXiv:2511.09677cs.LGstat.ML2025-11被引 4

通过分步训练提升生成模型探索能力,更好发现高奖励区域

Boosted GFlowNets: Improving Exploration via Sequential Learning

  • 逐次训练多个生成模型,每个优化前序模型未覆盖的残差奖励
  • 在多模态任务中样本多样性提升显著,高奖励区域覆盖率更高
  • 适合需要广泛探索的生成任务,如分子设计与复杂结构生成

生成流网络(GFlowNets)能按给定非负奖励比例采样组合对象,但实践中常陷入局部最优:易达区域的轨迹主导训练,难达模式因梯度消失导致学习信号弱,难以覆盖高奖励区域。本文提出增强型GFlowNets,通过顺序训练一组模型,每个模型优化前序模型已捕获分布的残差奖励。该残差机制重新激活低频区域的学习信号,在弱假设下保证单调不退化性——添加新模型不会恶化分布,通常还会有改进。实验表明,该方法在多模态合成基准和肽类设计任务中显著提升探索能力和样本多样性,同时保持标准轨迹平衡训练的稳定性和简洁性。

原文摘要 · Abstract (English)

Generative Flow Networks (GFlowNets) are powerful samplers for compositional objects that, by design, sample proportionally to a given non-negative reward. Nonetheless, in practice, they often struggle to explore the reward landscape evenly: trajectories toward easy-to-reach regions dominate training, while hard-to-reach modes receive vanishing or uninformative gradients, leading to poor coverage of high-reward areas. We address this imbalance with Boosted GFlowNets, a method that sequentially trains an ensemble of GFlowNets, each optimizing a residual reward that compensates for the mass already captured by previous models. This residual principle reactivates learning signals in underexplored regions and, under mild assumptions, ensures a monotone non-degradation property: adding boosters cannot worsen the learned distribution and typically improves it. Empirically, Boosted GFlowNets achieve substantially better exploration and sample diversity on multimodal synthetic benchmarks and peptide design tasks, while preserving the stability and simplicity of standard trajectory-balance training.

生成模型强化学习探索优化分子设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。