用进化方法评估动态博弈中联合策略的长期稳定性与优劣。
Ranking Joint Policies in Dynamic Games using Evolutionary Dynamics
- 将博弈转化为基于策略的实证形式,用α-排名分析长期动态。
- 在随机图着色任务中验证,识别出抗干扰且收益高的联合策略。
- 适合研究多智能体协作、博弈演化或策略评估的科研人员。
博弈论解概念(如纳什均衡)是寻找多智能体博弈中稳定联合行动的关键。然而研究表明,即使在仅有少数策略的简单双人博弈中,智能体交互的动态行为也难以达到纳什均衡,表现出复杂且不可预测的行为。相比之下,进化方法能描述策略的长期持续性并过滤瞬时策略,反映智能体交互的长期动态。本文目标是在动态博弈中识别出对变化具有抵抗性的稳定联合策略,并兼顾智能体收益。为此,基于前期成果,本文将动态博弈转换为基于策略而非行动的实证形式,应用进化方法α-排名来评估和排序策略组合,以揭示其长期动态表现。该方法不仅能识别通过长期交互表现出优势的联合策略,还提供透明可解释的评估框架。实验中,智能体需协同解决随机版本的图着色问题,将不同博弈风格视为策略,使用DQN训练实现这些策略的智能体政策,再通过模拟生成α-排名所需的收益矩阵,完成联合策略排序。
原文摘要 · Abstract (English)
Game-theoretic solution concepts, such as the Nash equilibrium, have been key to finding stable joint actions in multi-player games. However, it has been shown that the dynamics of agents' interactions, even in simple two-player games with few strategies, are incapable of reaching Nash equilibria, exhibiting complex and unpredictable behavior. Instead, evolutionary approaches can describe the long-term persistence of strategies and filter out transient ones, accounting for the long-term dynamics of agents' interactions. Our goal is to identify agents' joint strategies that result in stable behavior, being resistant to changes, while also accounting for agents' payoffs, in dynamic games. Towards this goal, and building on previous results, this paper proposes transforming dynamic games into their empirical forms by considering agents' strategies instead of agents' actions, and applying the evolutionary methodology $α$-Rank to evaluate and rank strategy profiles according to their long-term dynamics. This methodology not only allows us to identify joint strategies that are strong through agents' long-term interactions, but also provides a descriptive, transparent framework regarding the high ranking of these strategies. Experiments report on agents that aim to collaboratively solve a stochastic version of the graph coloring problem. We consider different styles of play as strategies to define the empirical game, and train policies realizing these strategies, using the DQN algorithm. Then we run simulations to generate the payoff matrix required by $α$-Rank to rank joint strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。