让网络化智能体在信息不全时仍能协作完成任务
Networked Agents in the Dark: Team Value Learning under Partial Observability
- 通过局部通信与共识机制实现团队价值函数近似
- 在部分可观测环境下性能优于已有方法
- 适合隐私敏感或通信不可靠的现实场景
我们提出一种新型协作多智能体强化学习方法,适用于网络化智能体。与依赖完整状态或联合观测的现有方法不同,本方法在部分可观测条件下让智能体学习达成共同目标。训练中,各智能体收集个体奖励,并通过局部通信近似团队价值函数,从而形成合作行为。为此,我们引入网络化动态部分可观测马尔可夫博弈框架,其中智能体在切换拓扑的通信网络上交互。所提出的分布式方法DNA-MARL采用共识机制进行局部通信,以梯度下降进行本地计算。该方法扩展了网络化智能体的应用范围,特别适用于具有隐私限制且消息可能无法送达的真实场景。我们在基准MARL场景中评估DNA-MARL,结果表明其性能显著优于先前方法。
原文摘要 · Abstract (English)
We propose a novel cooperative multi-agent reinforcement learning (MARL) approach for networked agents. In contrast to previous methods that rely on complete state information or joint observations, our agents must learn how to reach shared objectives under partial observability. During training, they collect individual rewards and approximate a team value function through local communication, resulting in cooperative behavior. To describe our problem, we introduce the networked dynamic partially observable Markov game framework, where agents communicate over a switching topology communication network. Our distributed method, DNA-MARL, uses a consensus mechanism for local communication and gradient descent for local computation. DNA-MARL increases the range of the possible applications of networked agents, being well-suited for real world domains that impose privacy and where the messages may not reach their recipients. We evaluate DNA-MARL across benchmark MARL scenarios. Our results highlight the superior performance of DNA-MARL over previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。