arXiv:2510.13343cs.MAcs.AI2025-10中稿 · as a short paper a…被引 1

提出考虑智能体行动顺序的Transformer多智能体强化学习模型,提升协同决策效率。

AOAD-MAT: Transformer-based multi-agent deep reinforcement learning model considering agents' order of action decisions

  • 引入基于Transformer的架构,动态调整智能体行动顺序以优化决策
  • 在星际争霸和多智能体MuJoCo上表现优于现有MAT及基线模型
  • 适合研究多智能体协同、动态调度或顺序敏感任务的科研人员

多智能体强化学习旨在训练共存于共享环境中的多个学习智能体的行为。近年来,如多智能体Transformer(MAT)和动作依赖深度Q-learning(ACE)等模型通过利用序列化决策过程显著提升了性能。然而,这些模型未显式考虑智能体决策顺序的重要性。本文提出一种考虑行动顺序的多智能体Transformer模型(AOAD-MAT),将行动顺序显式融入学习过程,使模型能够学习并预测最优的智能体行动顺序。该模型采用基于Transformer的演员-评论家架构,动态调整行动序列,并引入一个辅助子任务以预测下一个行动智能体,结合近端策略优化(PPO)损失函数,协同最大化序列决策优势。在星际争霸多智能体挑战赛(StarCraft Multi-Agent Challenge)和多智能体MuJoCo基准测试上的大量实验表明,所提方法优于现有的MAT及其他基线模型,验证了调整行动顺序在多智能体强化学习中的有效性。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning focuses on training the behaviors of multiple learning agents that coexist in a shared environment. Recently, MARL models, such as the Multi-Agent Transformer (MAT) and ACtion dEpendent deep Q-learning (ACE), have significantly improved performance by leveraging sequential decision-making processes. Although these models can enhance performance, they do not explicitly consider the importance of the order in which agents make decisions. In this paper, we propose an Agent Order of Action Decisions-MAT (AOAD-MAT), a novel MAT model that considers the order in which agents make decisions. The proposed model explicitly incorporates the sequence of action decisions into the learning process, allowing the model to learn and predict the optimal order of agent actions. The AOAD-MAT model leverages a Transformer-based actor-critic architecture that dynamically adjusts the sequence of agent actions. To achieve this, we introduce a novel MARL architecture that cooperates with a subtask focused on predicting the next agent to act, integrated into a Proximal Policy Optimization based loss function to synergistically maximize the advantage of the sequential decision-making. The proposed method was validated through extensive experiments on the StarCraft Multi-Agent Challenge and Multi-Agent MuJoCo benchmarks. The experimental results show that the proposed AOAD-MAT model outperforms existing MAT and other baseline models, demonstrating the effectiveness of adjusting the AOAD order in MARL.

多智能体Transformer强化学习决策顺序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。