提出新方法PANDA,解决多智能体竞争下的强化学习双层优化问题。
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games

- 基于极小极大博弈结构设计无超梯度的一阶优化方法
- 在无凸性假设下收敛至平稳点,样本复杂度为ε⁻³
- 适合激励设计等多方博弈场景的算法研究者
强化学习常具有分层结构:上层(UL)学习者选择模型参数,下层(LL)决策过程响应,自然形成双层优化问题。现有大多数双层强化学习方法假设下层为单策略马尔可夫决策过程,难以捕捉激励设计等应用中的竞争结构。本文研究下层为正则化极小极大零和马尔可夫博弈的双层优化问题,上层目标通过下层博弈的鞍点均衡进行优化。提出惩罚增强型尼卡伊多-伊索达下降-上升(PANDA)方法,基于尼卡伊多-伊索达函数的惩罚法,利用极小极大博弈结构避免计算上层超梯度且无需二阶信息。证明了PANDA在不依赖上层或下层目标凸性的条件下收敛至平稳点。此外,PANDA在$ ilde{igcal{O}}(ε^{-1})$次迭代内达到$ε$-平稳点,样本复杂度为$ ilde{igcal{O}}(ε^{-3})$,达到与单策略下层马尔可夫决策过程最优已知率相当的性能。实验表明PANDA显著优于相关基线方法。
原文摘要 · Abstract (English)
Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) decision-making process responds, naturally leading to a bilevel optimization problem. Most existing bilevel RL methods assume a single-policy LL Markov decision process (MDP), and therefore fail to capture competitive structures arising in applications such as incentive design, where multiple policies interact. We study bilevel optimization problems in which the LL problem is a regularized min-max zero-sum Markov game and the UL objective is optimized through the saddle-point equilibrium induced by the LL game. In this work, we propose penalty-augmented Nikaido-Isoda descent-ascent (PANDA), a penalty-based first-order policy-gradient method based on the Nikaido-Isoda function. By exploiting the min-max game structure, PANDA avoids computing UL hypergradients and does not require second-order information. We prove that PANDA converges to stationary points without convexity assumptions on either the UL or LL objectives. Moreover, PANDA reaches an $ε$-stationary point in $\tilde{\mathcal{O}}(ε^{-1})$ iterations with sample complexity $\tilde{\mathcal{O}}(ε^{-3})$, matching the best-known rates for bilevel RL with single-policy LL MDPs. Experiments demonstrate the superior performance of PANDA over closely related baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。