不确定性让博弈学习趋向极端策略,最终逼近纯纳什均衡。
The impact of uncertainty on regularized learning in games
- 引入随机扰动的FTRL学习动态,模拟真实环境中的观测误差。
- 无论噪声强弱,玩家策略在有限时间内必逼近任意纯策略邻域。
- 适合研究学习机制对博弈结果鲁棒性的影响,尤其关注零和博弈。
本文研究随机性和不确定性如何影响博弈中的学习过程。具体分析了受随机冲击影响的“跟随正则化领导者”(FTRL)动力学变体,其中玩家的收益观测与策略更新持续受到随机扰动。研究发现,在精确意义下,“不确定性有利于极端行为”:在任意博弈中,无论噪声水平如何,每位玩家的策略轨迹均在有限时间内进入任意小的纯策略邻域(我们给出了估计值)。即使未最终收敛至该策略,玩家仍会无限次地无限接近某个(可能不同)纯策略。这引出了在不确定性下哪些纯策略集可作为学习的稳健预测。我们证明:(a) FTRL在不确定性下的唯一可能极限是纯纳什均衡;(b) 纯策略集合稳定且吸引,当且仅当其对更好回应封闭。最后,针对确定性动力学具有循环性的博弈(如存在内部均衡的零和博弈),我们发现随机性会破坏这种循环,使随机动力学平均而言向边界漂移。
原文摘要 · Abstract (English)
In this paper, we investigate how randomness and uncertainty influence learning in games. Specifically, we examine a perturbed variant of the dynamics of "follow-the-regularized-leader" (FTRL), where the players' payoff observations and strategy updates are continually impacted by random shocks. Our findings reveal that, in a fairly precise sense, "uncertainty favors extremes": in any game, regardless of the noise level, every player's trajectory of play reaches an arbitrarily small neighborhood of a pure strategy in finite time (which we estimate). Moreover, even if the player does not ultimately settle at this strategy, they return arbitrarily close to some (possibly different) pure strategy infinitely often. This prompts the question of which sets of pure strategies emerge as robust predictions of learning under uncertainty. We show that (a) the only possible limits of the FTRL dynamics under uncertainty are pure Nash equilibria; and (b) a span of pure strategies is stable and attracting if and only if it is closed under better replies. Finally, we turn to games where the deterministic dynamics are recurrent - such as zero-sum games with interior equilibria - and we show that randomness disrupts this behavior, causing the stochastic dynamics to drift toward the boundary on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。