arXiv:2509.20408cs.LGcs.DC2025-09中稿 · NeurIPS被引 1

多智能体生成流网络理论框架,实现协同生成与奖励对齐。

A Theory of Multi-Agent Generative Flow Networks

  • 基于局部-全局原则训练多智能体生成流网络
  • 独立策略生成样本概率与奖励函数成正比
  • 适合需要协作生成的强化学习与采样任务

生成流网络利用流匹配损失学习随机策略,从一系列动作中生成对象,使得生成模式的概率与给定奖励成比例。然而,多智能体生成流网络(MA-GFlowNets)尚无理论框架。本文提出MA-GFlowNets的理论框架,支持多个智能体通过联合动作协同生成对象。进一步提出四种算法:集中式流网络用于集中训练,独立式流网络用于去中心化执行,联合式流网络实现集中训练+去中心化执行,及其条件更新版本。联合式训练基于局部-全局原则,将多个局部GFN训练为一个全局GFN,损失复杂度合理,并可借用传统GFN结果,提供理论保证:独立策略生成样本的概率与奖励函数成正比。实验表明,该框架优于强化学习和基于MCMC的方法。

原文摘要 · Abstract (English)

Generative flow networks utilize a flow-matching loss to learn a stochastic policy for generating objects from a sequence of actions, such that the probability of generating a pattern can be proportional to the corresponding given reward. However, a theoretical framework for multi-agent generative flow networks (MA-GFlowNets) has not yet been proposed. In this paper, we propose the theory framework of MA-GFlowNets, which can be applied to multiple agents to generate objects collaboratively through a series of joint actions. We further propose four algorithms: a centralized flow network for centralized training of MA-GFlowNets, an independent flow network for decentralized execution, a joint flow network for achieving centralized training with decentralized execution, and its updated conditional version. Joint Flow training is based on a local-global principle allowing to train a collection of (local) GFN as a unique (global) GFN. This principle provides a loss of reasonable complexity and allows to leverage usual results on GFN to provide theoretical guarantees that the independent policies generate samples with probability proportional to the reward function. Experimental results demonstrate the superiority of the proposed framework compared to reinforcement learning and MCMC-based methods.

多智能体生成模型流网络奖励对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。