提出新算法VMG,在不依赖奖励激励下实现高效多智能体强化学习。
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
- 通过偏向集体最优响应的模型估计引导探索
- 在线环境下近似最优后悔值,适用于零和与非零和博弈
- 支持独立更新,适合复杂函数逼近场景
多智能体强化学习(MARL)是多个智能体在共享未知环境中交互的核心框架。主流研究范式为马尔可夫博弈,目标是在样本高效条件下寻找纳什均衡(NE)或粗相关均衡(CCE)。现有高效方法要么需针对函数逼近设计特定不确定性估计,要么要求玩家间精细协调。本文提出一种新型基于模型的算法VMG,通过将经验模型参数偏向于固定其他玩家策略时具有更高集体最优响应值的方向,激励策略偏离当前均衡以促进探索。该算法对不同形式的函数逼近无感,支持所有玩家同时且解耦的策略更新。理论上,我们证明了在在线环境中,当使用线性函数逼近时,VMG能实现双人零和马尔可夫博弈中找寻纳什均衡、以及多人一般和马尔可夫博弈中找寻粗相关均衡的近似最优后悔率,几乎媲美需复杂不确定性量化的方法。
原文摘要 · Abstract (English)
Multi-agent reinforcement learning (MARL) lies at the heart of a plethora of applications involving the interaction of a group of agents in a shared unknown environment. A prominent framework for studying MARL is Markov games, with the goal of finding various notions of equilibria in a sample-efficient manner, such as the Nash equilibrium (NE) and the coarse correlated equilibrium (CCE). However, existing sample-efficient approaches either require tailored uncertainty estimation under function approximation, or careful coordination of the players. In this paper, we propose a novel model-based algorithm, called VMG, that incentivizes exploration via biasing the empirical estimate of the model parameters towards those with a higher collective best-response values of all the players when fixing the other players' policies, thus encouraging the policy to deviate from its current equilibrium for more exploration. VMG is oblivious to different forms of function approximation, and permits simultaneous and uncoupled policy updates of all players. Theoretically, we also establish that VMG achieves a near-optimal regret for finding both the NEs of two-player zero-sum Markov games and CCEs of multi-player general-sum Markov games under linear function approximation in an online environment, which nearly match their counterparts with sophisticated uncertainty quantification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。