arXiv:2608.05422cs.LGcs.MA2026-08中稿 · NeurIPS

将生成式采样框架拓展至不完全信息博弈,提升策略学习效率与性能。

IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

论文配图:IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
图 1 · 摘自论文原文
  • 提出IFlowNets框架,改进生成流网络在不完全信息博弈中的策略建模能力。
  • 在三个标准博弈环境中,性能与速度均优于或媲美OSMCCFR和传统强化学习方法。
  • 适用于需高效策略生成的博弈场景,尤其适合复杂不完全信息环境下的决策优化。

尽管许多算法结合强化学习(RL)与反事实遗憾(CFR)方法以平衡计算速度与性能,但针对不完全信息博弈中生成式采样框架的研究仍较少。本文将对抗性流网络(AFlowNets)扩展至不完全信息博弈,提出信息流网络(IFlowNets)。我们证明,先前在完全信息博弈中成立的生成流网络约束在不完全信息场景下会导致无效策略密度及不可行训练目标。所提方法有效缓解此问题,严格推广了AFlowNets。初步实验在三个标准博弈环境上显示,IFlowNets在性能与速度上均优于或媲美基于结果采样的蒙特卡洛反事实遗憾(OSMCCFR)及标准强化学习方法。

原文摘要 · Abstract (English)

While many algorithms blend reinforcement learning (RL) with counterfactual regret (CFR) methods to leverage tradeoffs in computational speed and performance, there are fewer investigations into generative sampling frameworks in game theoretic applications in incomplete information games. We extend a generative flow network framework, Adversarial Flow Networks (AFlowNets), to incomplete information games, called Information Flow Networks (IFNs). We prove that previously established constraints for generative flow networks in complete information games are inadmissible for obtaining valid densities (corresponding to player strategies) and a valid training objective. We show that our proposed generalization, IFlowNets, alleviates this issue and strictly generalizes AFlowNets. In preliminary results for three standard game environments, IFlowNets perform comparably to or better than Outcome Sampling Monte Carlo Counterfactual Regret (OSMCCFR) and standard RL-based methods in performance and speed.

博弈论生成模型强化学习不完全信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。