arXiv:2503.00372cs.MAcs.AI2025-03被引 5

用博弈论方法让智能体自动分组,提升协作效率。

Nucleolus Credit Assignment for Effective Coalitions in Multi-agent Reinforcement Learning

  • 基于核心值的信用分配,自动形成多个小团队。
  • 在复杂任务中实现更快学习和更高胜率,尤其在困难环境下。
  • 适合需要动态分工的多智能体协作场景。

在合作式多智能体强化学习中,传统方法通常依赖单一全体协作来完成复合任务,常导致性能不佳。本文提出一种基于合作博弈论核心值的信用分配机制,可自主将智能体划分为多个小型协作团体,有效识别并完成大任务中的子任务。所设计的核心值Q-learning能公平分配信用,核心值Q算子在理论上保证了学习收敛性和形成的小组稳定性。在追捕-猎物与星际争霸等不同难度场景下的实验表明,该方法在MARL训练过程中自发涌现出多个有效协作团体,显著提升了学习速度与表现,尤其在困难及超难环境中,胜率和累计奖励均优于四种基线方法。该核心值信用分配展现了在需高效子团队协作的复杂任务中的潜力。

原文摘要 · Abstract (English)

In cooperative multi-agent reinforcement learning (MARL), agents typically form a single grand coalition based on credit assignment to tackle a composite task, often resulting in suboptimal performance. This paper proposed a nucleolus-based credit assignment grounded in cooperative game theory, enabling the autonomous partitioning of agents into multiple small coalitions that can effectively identify and complete subtasks within a larger composite task. Specifically, our designed nucleolus Q-learning could assign fair credits to each agent, and the nucleolus Q-operator provides theoretical guarantees with interpretability for both learning convergence and the stability of the formed small coalitions. Through experiments on Predator-Prey and StarCraft scenarios across varying difficulty levels, our approach demonstrated the emergence of multiple effective coalitions during MARL training, leading to faster learning and superior performance in terms of win rate and cumulative rewards especially in hard and super-hard environments, compared to four baseline methods. Our nucleolus-based credit assignment showed the promise for complex composite tasks requiring effective subteams of agents.

多智能体博弈论协作优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。