用强化学习方法找到多个高分解空间,加速生成优质解。
Exploring Multiple High-Scoring Subspaces in Generative Flow Networks
- 引入组合多臂赌博机机制,筛选高质量动作
- 在多个高分空间中探索,提升解的质量与多样性
- 适用于复杂组合优化问题的高效求解
生成流网络(GFlowNets)作为概率采样框架,在通过逐步组合基本组件构建复杂组合对象方面展现出巨大潜力。然而,现有GFlowNets常因在广阔状态空间中过度探索,导致大量采样低奖励区域,并收敛至次优分布。有效引导其向高奖励解靠拢仍是挑战。本文提出CMAB-GFN,将组合多臂赌博机(CMAB)框架与GFlowNet策略结合,由CMAB组件剔除低质量动作,形成紧凑的高分解子空间供探索。将GFlowN限制在这些紧凑高分子空间中,可加速发现高价值候选解,同时跨子空间探索保障了多样性不丢失。在多个任务上的实验表明,CMAB-GFN生成的解奖励显著高于现有方法。
原文摘要 · Abstract (English)
As a probabilistic sampling framework, Generative Flow Networks (GFlowNets) show strong potential for constructing complex combinatorial objects through the sequential composition of elementary components. However, existing GFlowNets often suffer from excessive exploration over vast state spaces, leading to over-sampling of low-reward regions and convergence to suboptimal distributions. Effectively biasing GFlowNets toward high-reward solutions remains a non-trivial challenge. In this paper, we propose CMAB-GFN, which integrates a combinatorial multi-armed bandit (CMAB) framework with GFlowNet policies. The CMAB component prunes low-quality actions, yielding compact high-scoring subspaces for exploration. Restricting GFNs to these compact high-scoring subspaces accelerates the discovery of high-value candidates, while the exploration of different subspaces ensures that diversity is not sacrificed. Experimental results on multiple tasks demonstrate that CMAB-GFN generates higher-reward candidates than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。