arXiv:2511.11602cs.LGcs.GT2025-11被引 1

提出新型学习机制,让多智能体在噪声环境中更稳定收敛到最优策略。

Aspiration-based Perturbed Learning Automata in Games with Noisy Utility Measurements. Part A: Stochastic Stability in Non-zero-Sum Games

  • 基于期望值调整动作选择概率,结合重复尝试与满意度反馈。
  • 证明在非零和博弈中,系统可稳定收敛至纯纳什均衡。
  • 适用于分布式优化场景,尤其适合存在观测噪声的多智能体系统。

基于强化学习的方法在建模人类行为及工程优化中备受关注,尤其擅长过滤噪声观测。然而在分布式设置下,当各玩家独立运行学习动态时,难以保证收敛至理想的纯纳什均衡。现有研究仅限于势博弈和协调博弈等少数类型。本文提出一种新型基于收益的学习机制——基于期望的扰动学习自动机(APLA),其动作选择概率同时受重复选择频率与玩家满意度(期望)驱动。本文第一部分对带有噪声观测的多玩家正收益博弈中的APLA进行随机稳定性分析,首次将诱导的无限维马尔可夫链等价转化为有限维系统,从而实现对一般非零和博弈中随机稳定性的刻画。第二部分将进一步将其应用于弱循环博弈。

原文摘要 · Abstract (English)

Reinforcement-based learning has attracted considerable attention both in modeling human behavior as well as in engineering, for designing measurement- or payoff-based optimization schemes. Such learning schemes exhibit several advantages, especially in relation to filtering out noisy observations. However, they may exhibit several limitations when applied in a distributed setup. In multi-player weakly-acyclic games, and when each player applies an independent copy of the learning dynamics, convergence to (usually desirable) pure Nash equilibria cannot be guaranteed. Prior work has only focused on a small class of games, namely potential and coordination games. To address this main limitation, this paper introduces a novel payoff-based learning scheme for distributed optimization, namely aspiration-based perturbed learning automata (APLA). In this class of dynamics, and contrary to standard reinforcement-based learning schemes, each player's probability distribution for selecting actions is reinforced both by repeated selection and an aspiration factor that captures the player's satisfaction level. We provide a stochastic stability analysis of APLA in multi-player positive-utility games under the presence of noisy observations. This is the first part of the paper that characterizes stochastic stability in generic non-zero-sum games by establishing equivalence of the induced infinite-dimensional Markov chain with a finite dimensional one. In the second part, stochastic stability is further specialized to weakly acyclic games.

多智能体博弈学习稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。