arXiv:2608.08268math.OCcs.GT2026-08

在信息有限下,玩家仍能收敛到最优策略,适合研究博弈学习的场景。

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

  • 玩家仅观测自身历史与公共状态,用ε-贪婪算法自主学习。
  • 学习过程几乎必然收敛至完全信息下的纳什均衡,收敛速率可量化。
  • 公开市场总产量可加速收敛并改善整体福利,适合经济建模应用。

随着企业越来越多地使用机器学习进行战略决策,理解算法交互已成为运筹学与经济学的核心问题。本文研究在无限时域、非零和线性二次随机博弈中,当参与者信息结构极为简化(即对对手无知或策略性忽视),仅能观测共同状态及自身动作历史时的學習机制。我们分析一种异步去中心化学习过程:每位玩家独立运行单智能体ε-贪婪迭代最小二乘算法。尽管无法识别系统参数,我们证明玩家的学习动态几乎必然收敛至完全信息下的纳什均衡,并刻画了收敛速率。进一步将该框架应用于具有价格粘性的动态古诺竞争模型。数值实验验证了理论结果:在低与高价格粘性下,有限信息学习均导致企业利润下降;当价格粘性高时,总剩余减少,市场集中度上升。公开披露市场总产出可显著加速收敛并缓解福利损失。

原文摘要 · Abstract (English)

As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This paper studies learning in infinite-horizon, nonzero-sum linear-quadratic stochastic games under a radically uncoupled information structure, where players are either unaware of opponents or strategically oblivious, observing only a common state and their own action history. Under this minimal information, we analyze an asynchronous decentralized learning process in which each player independently runs a single-agent $ε$-greedy iterated least-squares algorithm. We prove that, despite being unable to identify the system parameters, players' learning dynamics converge almost surely to the complete-information Nash equilibrium and characterize the convergence rate. We then apply the framework to a dynamic Cournot competition with sticky prices. Numerical experiments validate the theoretical results and show that learning under limited information reduces firm profits under both low and high price stickiness, while total surplus declines and market concentration increases when price stickiness is high. Publicly revealing aggregate market output substantially accelerates convergence and mitigates these welfare losses.

博弈学习随机博弈动态竞争纳什均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。