arXiv:2506.11285cs.MAcs.GT2025-06被引 4

用博弈论方法解决多智能体临时协作中的贡献分配问题。

Shapley Machine: A Game-Theoretic Framework for N-Agent Ad Hoc Teamwork

  • 基于合作博弈理论构建动态协作框架,统一建模多智能体系统价值。
  • 提出类似λ-回报的值函数表示,实现对控制智能体的精确贡献评估。
  • 首次将博弈论中的谢林值直接融入强化学习,适合研究智能体协作机制的研究者。

开放多智能体系统在智能电网、群体机器人等现实应用中日益重要。本文研究一种新提出的开放多智能体系统问题——n-agent临时协作(NAHT),其中仅部分智能体受控。现有方法多依赖启发式设计,缺乏理论依据且贡献分配模糊。为此,本文从合作博弈论视角建模和求解NAHT:首先将开放系统的价值视为一组基博弈生成的合作博弈空间中的实例;进而扩展该空间与状态空间以适应动态场景,从而刻画NAHT。通过合理假设基博弈值对应于不同时间跨度的n步回报,将NAHT的状态值表示为类似λ-回报的形式。进一步,利用谢林值对受控智能体的贡献进行信用分配。不同于传统显式构造谢林值的方法,本文通过满足唯一定义其的三个公理,在扩展的博弈空间上精准构造谢林值。为在动态场景中估计谢林值,提出一种类似TD(λ)的算法,所获强化学习算法称为Shapley Machine。据我们所知,这是首次将合作博弈论概念直接关联到强化学习概念。实验验证了Shapley Machine的有效性,并证明理论合理性。

原文摘要 · Abstract (English)

Open multi-agent systems are increasingly important in modeling real-world applications, such as smart grids, swarm robotics, etc. In this paper, we aim to investigate a recently proposed problem for open multi-agent systems, referred to as n-agent ad hoc teamwork (NAHT), where only a number of agents are controlled. Existing methods tend to be based on heuristic design and consequently lack theoretical rigor and ambiguous credit assignment among agents. To address these limitations, we model and solve NAHT through the lens of cooperative game theory. More specifically, we first model an open multi-agent system, characterized by its value, as an instance situated in a space of cooperative games, generated by a set of basis games. We then extend this space, along with the state space, to accommodate dynamic scenarios, thereby characterizing NAHT. Exploiting the justifiable assumption that basis game values correspond to a sequence of n-step returns with different horizons, we represent the state values for NAHT in a form similar to $λ$-returns. Furthermore, we derive Shapley values to allocate state values to the controlled agents, as credits for their contributions to the ad hoc team. Different from the conventional approach to shaping Shapley values in an explicit form, we shape Shapley values by fulfilling the three axioms uniquely describing them, well defined on the extended game space describing NAHT. To estimate Shapley values in dynamic scenarios, we propose a TD($λ$)-like algorithm. The resulting reinforcement learning (RL) algorithm is referred to as Shapley Machine. To our best knowledge, this is the first time that the concepts from cooperative game theory are directly related to RL concepts. In experiments, we demonstrate the effectiveness of Shapley Machine and verify reasonableness of our theory.

多智能体博弈论强化学习贡献分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。