首个证明深度神经网络在分布式多智能体强化学习中有限时间全局收敛的理论工作
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning
- 采用非线性深度神经网络构建演员-评论家框架,实现去中心化多智能体协同决策
- 理论证明算法在总迭代次数T下达到O(1/T)的有限时间全局最优收敛速度
- 填补了深度神经网络在多智能体强化学习中理论与实践的巨大鸿沟,适合研究者参考
去中心化多智能体强化学习(MARL)中的演员-评论家方法可在无需集中协调的情况下实现协作最优决策,广泛应用于实际场景。然而,现有理论研究大多局限于线性函数近似下的平稳解保证,导致深度神经网络在实践中广泛应用与当前理论理解之间存在显著差距。本文首次提出一种基于深度神经网络的去中心化MARL演员-评论家方法,其中演员和评论家均具有内在非线性。我们证明该方法在总迭代次数为T时,具备全局最优性保障和O(1/T)的有限时间收敛速率。这是MARL领域中首个针对深度神经网络演员-评论家方法的全局收敛结果。我们还进行了大量数值实验,验证了理论结论。
原文摘要 · Abstract (English)
Actor-critic methods for decentralized multi-agent reinforcement learning (MARL) facilitate collaborative optimal decision making without centralized coordination, thus enabling a wide range of applications in practice. To date, however, most theoretical convergence studies for existing actor-critic decentralized MARL methods are limited to the guarantee of a stationary solution under the linear function approximation. This leaves a significant gap between the highly successful use of deep neural actor-critic for decentralized MARL in practice and the current theoretical understanding. To bridge this gap, in this paper, we make the first attempt to develop a deep neural actor-critic method for decentralized MARL, where both the actor and critic components are inherently non-linear. We show that our proposed method enjoys a global optimality guarantee with a finite-time convergence rate of O(1/T), where T is the total iteration times. This marks the first global convergence result for deep neural actor-critic methods in the MARL literature. We also conduct extensive numerical experiments, which verify our theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。