arXiv:2507.17183cs.GTcs.AI2025-07被引 3

研究大规模异质智能体如何通过悔恨匹配达成共识与均衡。

Regret Minimization in Population Network Games: Vanishing Heterogeneity and Convergence to Equilibria

  • 用后悔分布的连续性方程分析智能体行为演化。
  • 发现后悔方差随时间减小,异质性消失并趋于一致。
  • 适用于竞争与合作场景,为均衡选择提供新视角。

理解大规模多智能体在博弈中的行为仍是多智能体系统的核心挑战。本文通过分析平滑后悔匹配如何使大量初始策略各异的异质智能体趋向统一行为,揭示了异质性在均衡形成中的作用。将系统状态建模为后悔的概率分布,并通过连续性方程分析其演化,发现一个普遍现象:后悔分布的方差随时间减小,导致异质性消失,智能体间出现共识。这一结果可证明在竞争与合作的多智能体场景中均收敛至量化响应均衡。本研究深化了对多智能体学习的理论理解,并为多样博弈场景下的均衡选择提供了新视角。

原文摘要 · Abstract (English)

Understanding and predicting the behavior of large-scale multi-agents in games remains a fundamental challenge in multi-agent systems. This paper examines the role of heterogeneity in equilibrium formation by analyzing how smooth regret-matching drives a large number of heterogeneous agents with diverse initial policies toward unified behavior. By modeling the system state as a probability distribution of regrets and analyzing its evolution through the continuity equation, we uncover a key phenomenon in diverse multi-agent settings: the variance of the regret distribution diminishes over time, leading to the disappearance of heterogeneity and the emergence of consensus among agents. This universal result enables us to prove convergence to quantal response equilibria in both competitive and cooperative multi-agent settings. Our work advances the theoretical understanding of multi-agent learning and offers a novel perspective on equilibrium selection in diverse game-theoretic scenarios.

多智能体博弈论均衡收敛后悔匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。