arXiv:2506.04265cs.MAcs.AI2025-06被引 1

用博弈论核心思想分配奖励,让多智能体协作更高效。

Cooperative Game-Theoretic Credit Assignment for Multi-Agent Policy Gradients via the Core

  • 基于合作博弈的核心分配机制,评估不同组合的贡献。
  • 在多种任务上超越基线,提升智能体协同表现。
  • 适合研究多智能体强化学习与奖励设计的学者。

本文聚焦于合作式多智能体强化学习中的信用分配问题。传统全局优势共享方法难以捕捉不同智能体的联合贡献,导致策略优化不足。为此,我们从联盟视角重新审视策略更新过程,提出基于合作博弈论核心分配的CORA方法。通过评估不同联盟的边际贡献,并结合截断双Q学习以缓解高估偏差,CORA估计联盟级优势。核心公式对分配信用施加联盟级下界约束,使高贡献联盟的参与智能体获得更强激励,从而将全局优势归因于不同联盟策略,促进协调最优行为。为降低计算开销,采用随机联盟采样近似核心分配。在矩阵博弈、微分博弈及多智能体协作基准上的实验表明,该方法显著优于基线。结果凸显了联盟级信用分配与合作博弈在推进多智能体学习中的重要性。

原文摘要 · Abstract (English)

This work focuses on the credit assignment problem in cooperative multi-agent reinforcement learning (MARL). Sharing the global advantage among agents often leads to insufficient policy optimization, as it fails to capture the coalitional contributions of different agents. In this work, we revisit the policy update process from a coalitional perspective and propose CORA, an advantage allocation method guided by a cooperative game-theoretic core allocation. By evaluating the marginal contributions of different coalitions and combining clipped double Q-learning to mitigate overestimation bias, CORA estimates coalition-wise advantages. The core formulation enforces coalition-wise lower bounds on allocated credits, so that coalitions with higher advantages receive stronger total incentives for their participating agents, enabling the global advantage to be attributed to different coalition strategies and promoting coordinated optimal behavior. To reduce computational overhead, we employ random coalition sampling to approximate the core allocation efficiently. Experiments on matrix games, differential games, and multi-agent collaboration benchmarks demonstrate that our method outperforms baselines. These findings highlight the importance of coalition-level credit assignment and cooperative games for advancing multi-agent learning.

多智能体信用分配博弈论强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。