提出可扩展的集中式策略框架,提升多智能体强化学习性能
Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning
- 采用排列等变网络实现全集中式训练与执行
- 在MPE/SMAC/RWARE上显著超越标准CTDE算法
- 轻量且易实现,适合大规模多智能体场景
集中式训练与去中心化执行(CTDE)范式在多智能体强化学习中广受关注,是众多近期算法的基础。然而,去中心化策略受限于部分可观测性,性能通常不如集中式策略;而完全集中式方法在智能体数量增加时面临可扩展性挑战。本文提出集中式排列等变(CPE)学习,一种采用全集中式策略的训练与执行框架。该方法利用新型排列等变架构——全局-局部排列等变(GLPE)网络,具备轻量、可扩展、易实现的优点。实验表明,CPE可无缝集成至值分解与演员-评论家方法,在MPE、SMAC和RWARE等合作基准任务上显著提升标准CTDE算法性能,并达到当前最优RWARE实现的水平。
原文摘要 · Abstract (English)
The Centralized Training with Decentralized Execution (CTDE) paradigm has gained significant attention in multi-agent reinforcement learning (MARL) and is the foundation of many recent algorithms. However, decentralized policies operate under partial observability and often yield suboptimal performance compared to centralized policies, while fully centralized approaches typically face scalability challenges as the number of agents increases. We propose Centralized Permutation Equivariant (CPE) learning, a centralized training and execution framework that employs a fully centralized policy to overcome these limitations. Our approach leverages a novel permutation equivariant architecture, Global-Local Permutation Equivariant (GLPE) networks, that is lightweight, scalable, and easy to implement. Experiments show that CPE integrates seamlessly with both value decomposition and actor-critic methods, substantially improving the performance of standard CTDE algorithms across cooperative benchmarks including MPE, SMAC, and RWARE, and matching the performance of state-of-the-art RWARE implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。