通过分层领导者机制,让多智能体协作更高效可靠。
Hierarchical Lead Critic based Multi-Agent Reinforcement Learning
- 设计分层领导批评者架构,融合高低层级视角
- 在多个基准上实现更高性能与样本效率
- 适合大规模、部分可观测的协作场景
合作式多智能体强化学习(MARL)能解决需协同完成的复杂任务,但通常局限于局部(独立学习)或全局(集中学习)视角。本文提出一种新型分层训练方案与MARL架构,从不同层次的多个视角进行学习。我们设计了分层领导批评者(HLC),受团队结构中自然涌现的上下级关系启发,高层目标引导与底层执行相结合。实验表明,引入多层级结构并融合局部与全局视角,可显著提升性能,具备高样本效率和鲁棒策略。在合作性、非通信及部分可观测的MARL基准测试中,HLC优于单一层级基线,并随智能体数量和任务难度增加仍保持良好扩展性。
原文摘要 · Abstract (English)
Cooperative Multi-Agent Reinforcement Learning (MARL) solves complex tasks that require coordination from multiple agents, but is often limited to either local (independent learning) or global (centralized learning) perspectives. In this paper, we introduce a novel sequential training scheme and MARL architecture, which learns from multiple perspectives on different hierarchy levels. We propose the Hierarchical Lead Critic (HLC) - inspired by natural emerging distributions in team structures, where following high-level objectives combines with low-level execution. HLC demonstrates that introducing multiple hierarchies, leveraging local and global perspectives, can lead to improved performance with high sample efficiency and robust policies. Experimental results conducted on cooperative, non-communicative, and partially observable MARL benchmarks demonstrate that HLC outperforms single hierarchy baselines and scales robustly with increasing amounts of agents and difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。