降低去中心化多智能体强化学习中通信带来的学习方差
Reducing Variance Caused by Communication in Decentralized Multi-agent Deep Reinforcement Learning
- 提出模块化方法减少策略梯度中的通信方差
- 在星雀和交通路口任务中提升性能并降低训练方差
- 适合研究去中心化多智能体系统与通信优化的读者
在去中心化多智能体深度强化学习(MADRL)中,通信有助于智能体更好地理解环境并协调行为。然而,通信可能引入不确定性,从而增加学习过程中的方差。本文针对一种具有通信机制的去中心化MADRL设置,进行理论分析,研究通信导致的策略梯度方差。提出模块化技术以降低训练期间的策略梯度方差。将这些技术应用于两个现有的去中心化MADRL通信算法,并在星雀多智能体挑战(StarCraft Multi-Agent Challenge)和交通路口(Traffic Junction)领域多个任务上进行评估。结果表明,结合所提技术的通信方法不仅实现了高性能智能体,还在训练过程中显著降低了策略梯度方差。
原文摘要 · Abstract (English)
In decentralized multi-agent deep reinforcement learning (MADRL), communication can help agents to gain a better understanding of the environment to better coordinate their behaviors. Nevertheless, communication may involve uncertainty, which potentially introduces variance to the learning of decentralized agents. In this paper, we focus on a specific decentralized MADRL setting with communication and conduct a theoretical analysis to study the variance that is caused by communication in policy gradients. We propose modular techniques to reduce the variance in policy gradients during training. We adopt our modular techniques into two existing algorithms for decentralized MADRL with communication and evaluate them on multiple tasks in the StarCraft Multi-Agent Challenge and Traffic Junction domains. The results show that decentralized MADRL communication methods extended with our proposed techniques not only achieve high-performing agents but also reduce variance in policy gradients during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。