提出新方法解决多人随机微分博弈中的均衡学习问题
Learning Distributed Equilibria in Linear-Quadratic Stochastic Differential Games: An $α$-Potential Approach
- 基于α-势函数构建分布式策略更新机制
- 对称情况下全局线性收敛,复杂度随规模线性增长
- 适用于大规模博弈,尤其适合异步独立学习场景
本文研究N人线性-二次随机微分博弈中独立策略梯度(PG)学习的性质。每位玩家仅根据自身状态采用分布式策略,并独立使用自身目标函数梯度进行更新。通过证明该LQ博弈具有α-势结构(α由成对交互不对称程度决定),建立了方法的全局线性收敛性。在成对对称交互下,构造了仿射分布式均衡,且独立PG方法全局收敛,复杂度随群体规模线性增长,与期望精度呈对数关系。在非对称交互下,独立投影PG算法线性收敛至近似均衡,次优性与不对称程度成正比。数值实验验证了理论结果在对称与非对称交互网络中的有效性。
原文摘要 · Abstract (English)
We analyze independent policy-gradient (PG) learning in $N$-player linear-quadratic (LQ) stochastic differential games. Each player employs a distributed policy that depends only on its own state and updates the policy independently using the gradient of its own objective. We establish global linear convergence of these methods to an equilibrium by showing that the LQ game admits an $α$-potential structure, with $α$ determined by the degree of pairwise interaction asymmetry. For pairwise-symmetric interactions, we construct an affine distributed equilibrium by minimizing the potential function and show that independent PG methods converge globally to this equilibrium, with complexity scaling linearly in the population size and logarithmically in the desired accuracy. For asymmetric interactions, we prove that independent projected PG algorithms converge linearly to an approximate equilibrium, with suboptimality proportional to the degree of asymmetry. Numerical experiments confirm the theoretical results across both symmetric and asymmetric interaction networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。