提出无需模型、无需梯度的博弈学习框架,实现零和与同利博弈的收敛求解。
Actor-Dual-Critic Dynamics for Zero-sum and Identical-Interest Stochastic Games
- 基于双评分子架构,快评子响应实时收益,慢评子逼近动态规划解。
- 在无限时域下证明了零和与同利博弈的近似均衡收敛性。
- 首个无模型、全去中心化且带理论保证的收益驱动学习算法。
我们提出一种新型的独立式、基于收益的学习框架,适用于随机博弈,具有无模型、游戏无关、无梯度的特点。学习动态采用类最优响应的演员-评论家结构,代理通过两个不同评论家的反馈更新策略:快速评论家在信息受限下直观响应观测到的收益,慢速评论家则逐步逼近底层动态规划问题的解。关键在于,学习过程依赖于对观测收益的平滑最优响应进行非均衡适应。我们在两代理零和博弈和多代理同利博弈的无限时域场景中,建立了向(近似)均衡收敛的理论结果。这为两类博弈提供了少数几个具备理论保证的基于收益、完全去中心化的学习算法之一。实验结果进一步验证了该方法在两类博弈中的鲁棒性和有效性。
原文摘要 · Abstract (English)
We propose a novel independent and payoff-based learning framework for stochastic games that is model-free, game-agnostic, and gradient-free. The learning dynamics follow a best-response-type actor-critic architecture, where agents update their strategies (actors) using feedback from two distinct critics: a fast critic that intuitively responds to observed payoffs under limited information, and a slow critic that deliberatively approximates the solution to the underlying dynamic programming problem. Crucially, the learning process relies on non-equilibrium adaptation through smoothed best responses to observed payoffs. We establish convergence to (approximate) equilibria in two-agent zero-sum and multi-agent identical-interest stochastic games over an infinite horizon. This provides one of the first payoff-based and fully decentralized learning algorithms with theoretical guarantees in both settings. Empirical results further validate the robustness and effectiveness of the proposed approach across both classes of games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。