用注意力机制提升多智能体协作效率,让团队配合更默契。
Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
- 引入多头注意力的演员-评论家框架,实现动态跨智能体通信。
- 在模拟足球环境中胜率、进球差等指标全面超越基准方法。
- 适合研究多智能体协作、强化学习与角色分工的学者参考。
本文提出团队注意力演员-评论家(TAAC)算法,用于增强合作环境下的多智能体协作。该算法采用集中训练/集中执行范式,在演员和评论家中均引入多头注意力机制,实现智能体间动态通信,使智能体可主动查询队友,有效应对联合动作空间的指数级增长,同时保障高程度协作。我们还设计了惩罚性损失函数,促进智能体间形成多样化且互补的角色分工。在模拟足球环境中,TAAC 与包括近端策略优化和多智能体演员注意力评论家在内的基准算法进行对比。实验结果表明,TAAC 在胜率、进球差、Elo评分、智能体间连通性、空间分布均衡性以及传球交接等战术互动频率等多个指标上表现更优。
原文摘要 · Abstract (English)
This paper introduces Team-Attention-Actor-Critic (TAAC), a reinforcement learning algorithm designed to enhance multi-agent collaboration in cooperative environments. TAAC employs a Centralized Training/Centralized Execution scheme incorporating multi-headed attention mechanisms in both the actor and critic. This design facilitates dynamic, inter-agent communication, allowing agents to explicitly query teammates, thereby efficiently managing the exponential growth of joint-action spaces while ensuring a high degree of collaboration. We further introduce a penalized loss function which promotes diverse yet complementary roles among agents. We evaluate TAAC in a simulated soccer environment against benchmark algorithms representing other multi-agent paradigms, including Proximal Policy Optimization and Multi-Agent Actor-Attention-Critic. We find that TAAC exhibits superior performance and enhanced collaborative behaviors across a variety of metrics (win rates, goal differentials, Elo ratings, inter-agent connectivity, balanced spatial distributions, and frequent tactical interactions such as ball possession swaps).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。