arXiv:2512.08132cs.GTcs.LG2025-12NeurIPS被引 1

研究多智能体学习在不确定性下的长期行为,发现均衡附近有稳定分布。

Multi-agent learning under uncertainty: Recurrence vs. concentration

  • 用连续与离散时间模型分析正则化学习动态
  • 强单调博弈中策略分布高度集中在均衡附近
  • 非强单调时性能崩溃,揭示算法局限性

本文研究不确定性环境下多智能体学习的收敛特性。针对连续博弈中的两种正则化学习随机模型(连续时间与离散时间),我们刻画了其长期行为。与确定性或学习率趋零的情形不同,一般情况下动态不收敛。取而代之的是,我们关注长期中哪些动作被更频繁地执行及其频率差异。在强单调博弈中,学习动态虽可能无限次偏离均衡,但总能在有限时间内返回其邻域,且长期分布高度集中于该邻域。我们量化了这种集中程度,并证明若博弈非强单调,则这些有利性质均会失效,凸显了持续随机性下正则化学习的边界。

原文摘要 · Abstract (English)

In this paper, we examine the convergence landscape of multi-agent learning under uncertainty. Specifically, we analyze two stochastic models of regularized learning in continuous games -- one in continuous and one in discrete time with the aim of characterizing the long-run behavior of the induced sequence of play. In stark contrast to deterministic, full-information models of learning (or models with a vanishing learning rate), we show that the resulting dynamics do not converge in general. In lieu of this, we ask instead which actions are played more often in the long run, and by how much. We show that, in strongly monotone games, the dynamics of regularized learning may wander away from equilibrium infinitely often, but they always return to its vicinity in finite time (which we estimate), and their long-run distribution is sharply concentrated around a neighborhood thereof. We quantify the degree of this concentration, and we show that these favorable properties may all break down if the underlying game is not strongly monotone -- underscoring in this way the limits of regularized learning in the presence of persistent randomness and uncertainty.

多智能体学习动态博弈论随机系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。