用内部状态提升部分可观测博弈中的多智能体学习效果
Internal State-Based Policy Gradient Methods for Partially Observable Markov Potential Games
- 引入内部状态压缩信息,避免维度爆炸
- 理论证明方法在有限步内收敛,误差可分解为统计与近似项
- 实验证明相比仅用当前观测性能显著提升
本文研究部分可观测马尔可夫潜在博弈中的多智能体强化学习。由于部分可观测性、分布式信息和维数灾难,求解极具挑战。首先,通过公共信息框架,使智能体基于共享与局部信息行动;其次,引入内部状态以压缩累积信息,防止其随时间无界增长。进而提出基于内部状态的自然策略梯度方法,用于寻找马尔可夫潜在博弈的纳什均衡。主要贡献是建立了该方法的非渐近收敛边界,其分解为两类可解释成分:标准潜在博弈中出现的统计误差项,以及由有限状态控制器引入的近似误差。多个部分可观测环境下的仿真表明,使用有限状态控制器的方法在性能上持续优于仅依赖当前观测的设定。
原文摘要 · Abstract (English)
This letter studies multi-agent reinforcement learning in partially observable Markov potential games. Solving this problem is challenging due to partial observability, decentralized information, and the curse of dimensionality. First, to address the first two challenges, we leverage the common information framework, which allows agents to act based on both shared and local information. Second, to ensure tractability, we study an internal state that compresses accumulated information, preventing it from growing unboundedly over time. We then implement an internal state-based natural policy gradient method to find Nash equilibria of the Markov potential game. Our main contribution is to establish a non-asymptotic convergence bound for this method. Our theoretical bound decomposes into two interpretable components: a statistical error term that also arises in standard Markov potential games, and an approximation error capturing the use of finite-state controllers. Finally, simulations across multiple partially observable environments demonstrate that the proposed method using finite-state controllers achieves consistent improvements in performance compared to the setting where only the current observation is used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。