在通信受限下,用采样均值场强化学习实现多智能体近似纳什均衡。
Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
- 全局智能体仅观测部分局部状态,通过交替学习更新策略。
- 收敛至误差为 $\widetilde{O}(1/\sqrt{k})$ 的近似纳什均衡。
- 适用于大规模机器人协同控制等高维通信受限场景。
许多大规模平台和网络化控制系统具有一个中心决策者,与大量智能体在严格可观测性约束下交互。针对此类应用,我们研究了一个协作马尔可夫博弈,包含一个全局智能体和 $n$ 个同质局部智能体,在通信受限环境下,全局智能体每步仅能观测 $k$ 个局部智能体的状态。本文提出一种交替学习框架(ALTERNATING-MARL),其中全局智能体基于固定局部策略执行子采样的均值场 $Q$-学习,局部智能体则在诱导的马尔可夫决策过程(MDP)中优化自身策略。理论证明,该近似最优响应动态收敛至 $\widetilde{O}(1/\sqrt{k})$-近似纳什均衡,并分离了联合状态空间与动作空间的样本复杂度。数值仿真验证了其在多机器人控制中的有效性。
原文摘要 · Abstract (English)
Many large-scale platforms and networked control systems have a centralized decision maker interacting with a massive population of agents under strict observability constraints. Motivated by such applications, we study a cooperative Markov game with a global agent and $n$ homogeneous local agents in a communication-constrained regime, where the global agent only observes a subset of $k$ local agent states per time step. We propose an alternating learning framework $(\texttt{ALTERNATING-MARL})$, where the global agent performs subsampled mean-field $Q$-learning against a fixed local policy, and local agents update by optimizing in an induced MDP. We prove that these approximate best-response dynamics converge to an $\widetilde{O}(1/\sqrt{k})$-approximate Nash Equilibrium, while separating the sample complexities between the joint state and action spaces. Finally, we validate our results in numerical simulations for multi-robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。