提出新算法求解信息不对称的博弈,无需双方完全知情也能快速收敛。
Asymmetric Nash Seeking via Best Response Maps: Global Linear Convergence and Robustness to Inexact Reaction Models
- 设计异步梯度-响应迭代,仅需一方提供最优反应映射。
- 证明全局线性收敛,误差受估计偏差控制在O(ε)内。
- 适合实际中难以获取完整对手信息的多智能体系统设计。
纳什均衡为多智能体决策与控制中的交互提供了严谨建模框架。然而,许多均衡求解方法隐含假设各智能体可获知其他方的目标与约束,这在实践中往往不成立。本文研究一类具有解耦可行集的双人非对称信息约束博弈:玩家1掌握自身目标与约束,玩家2仅通过最优反应映射可得。针对此类博弈,提出一种异步投影梯度下降-最优反应迭代算法,无需双方完全知晓对方优化问题。在适当正则条件下,证明了纳什均衡的存在唯一性,并在最优反应映射精确时建立了算法的全局线性收敛性。考虑到最优反应映射常为学习或估计所得,进一步分析了近似情形:当逼近误差一致有界于ε时,迭代序列进入一个明确的O(ε)邻域内。基准博弈的数值结果验证了预测的收敛行为与误差标度关系。
原文摘要 · Abstract (English)
Nash equilibria provide a principled framework for modeling interactions in multi-agent decision-making and control. However, many equilibrium-seeking methods implicitly assume that each agent has access to the other agents' objectives and constraints, an assumption that is often unrealistic in practice. This letter studies a class of asymmetric-information two-player constrained games with decoupled feasible sets, in which Player 1 knows its own objective and constraints while Player 2 is available only through a best-response map. For this class of games, we propose an asymmetric projected gradient descent-best response iteration that does not require full mutual knowledge of both players' optimization problems. Under suitable regularity conditions, we establish the existence and uniqueness of the Nash equilibrium and prove global linear convergence of the proposed iteration when the best-response map is exact. Recognizing that best-response maps are often learned or estimated, we further analyze the inexact case and show that, when the approximation error is uniformly bounded by $\varepsilon$, the iterates enter an explicit $O(\varepsilon)$ neighborhood of the true Nash equilibrium. Numerical results on a benchmark game corroborate the predicted convergence behavior and error scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。