通过分层指挥官机制,提升多智能体强化学习的协作效率与探索能力。
HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement Learning
- 引入分层指挥官框架,协调各智能体策略更新方向。
- 在三个基准测试中显著优于主流MARL方法,提升合作效率与稳定性。
- 训练时集中优化,执行时无需通信,适合复杂协同任务。
在合作式多智能体强化学习中,高效探索对优化联合策略性能至关重要。然而,现有方法通常通过独立智能体探索来更新联合策略,缺乏智能体间的协调,从而限制了联合策略的表达能力与探索范围。为此,我们提出一种基于指挥官的联合策略框架,直接增强联合策略的表达能力并实现探索协调。进一步地,我们设计了分层指挥官策略优化(HCPO)算法,引导指挥官与智能体的策略更新朝向性能提升方向进行。理论分析证明了联合策略优化过程的单调性。通过部署局部指挥官,HCPO在保留集中式训练优势的同时,避免了执行阶段的智能体间通信。我们在三个挑战性基准测试——星雀II多智能体挑战、多智能体MuJoCo和多智能体粒子环境上评估了HCPO,结果表明其在合作效率与稳定性方面均优于现有的先进MARL基线方法。
原文摘要 · Abstract (English)
In cooperative Multi-Agent Reinforcement Learning (MARL), efficient exploration is crucial for optimizing the performance of joint policy. However, existing methods often update joint policies via independent agent exploration, without coordination among agents, which inherently constrains the expressive capacity and exploration of joint policies. To address this issue, we propose a conductor-based joint policy framework that directly enhances the expressive capacity of joint policies and coordinates exploration. In addition, we develop a Hierarchical Conductor-based Policy Optimization (HCPO) algorithm that instructs policy updates for the conductor and agents in a direction aligned with performance improvement. A rigorous theoretical guarantee further establishes the monotonicity of the joint policy optimization process. By deploying local conductors, HCPO retains centralized training benefits while eliminating inter-agent communication during execution. Finally, we evaluate HCPO on three challenging benchmarks: StarCraftII Multi-agent Challenge, Multi-agent MuJoCo, and Multi-agent Particle Environment. The results indicate that HCPO outperforms competitive MARL baselines regarding cooperative efficiency and stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。