用可学习的信念状态提升部分可观测多智能体强化学习性能
Belief States for Cooperative Multi-Agent Reinforcement Learning under Partial Observability
- 通过自监督预训练概率信念模型,估计系统真实状态
- 分离信念与强化学习任务,加速收敛并提升最终表现
- 支持完全去中心化训练执行,适合复杂协作场景
部分可观测环境中的强化学习通常具有挑战性,因需代理学习对底层系统状态的估计。这一挑战在多智能体环境中进一步加剧,因为代理同时学习并影响底层状态及彼此观测。本文提出使用对系统底层状态的可学习信念来克服这些挑战,实现完全去中心化的训练与执行。该方法利用状态信息以自监督方式预训练概率信念模型,生成的信念状态既包含推断出的状态信息,也包含对这些信息的不确定性。随后将这些信念状态用于基于状态的强化学习算法,构建端到端的协作式多智能体强化学习框架。通过分离信念学习与强化学习任务,显著简化了策略与价值函数的学习过程,提升了收敛速度与最终性能。我们在多种设计用于体现不同部分可观测性变体的多智能体任务上评估了该方法。
原文摘要 · Abstract (English)
Reinforcement learning in partially observable environments is typically challenging, as it requires agents to learn an estimate of the underlying system state. These challenges are exacerbated in multi-agent settings, where agents learn simultaneously and influence the underlying state as well as each others' observations. We propose the use of learned beliefs on the underlying state of the system to overcome these challenges and enable reinforcement learning with fully decentralized training and execution. Our approach leverages state information to pre-train a probabilistic belief model in a self-supervised fashion. The resulting belief states, which capture both inferred state information as well as uncertainty over this information, are then used in a state-based reinforcement learning algorithm to create an end-to-end model for cooperative multi-agent reinforcement learning under partial observability. By separating the belief and reinforcement learning tasks, we are able to significantly simplify the policy and value function learning tasks and improve both the convergence speed and the final performance. We evaluate our proposed method on diverse partially observable multi-agent tasks designed to exhibit different variants of partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。