研究游戏偏好如何决定多人在线学习的长期行为,发现偏好不总能预测稳定结果。
What preferences can - and cannot - predict in multi-agent online learning
- 用偏好图分析学习动态的稳定性,发现偏好稳定是动态稳定的必要条件。
- 构造三玩家博弈反例,证明偏好稳定集可能动态不稳定。
- 提出可验证的收益条件‘聚合偏离鲁棒性’,确保任意纯策略组合的长期稳定。
我们研究博弈中序数偏好概念与学习动态长期行为之间的关系,特别关注博弈的偏好图能否决定无悔学习动态(如跟随正则化领导者,FTRL)的结果。一方面,我们证明每个动态稳定集的骨架(即包含的纯策略组合)必须具有偏好稳定性——对有利偏离封闭。进一步探讨逆问题:何时偏好能确定长期行为?我们发现,在子博弈(通过限制玩家行动集得到的纯策略子集)情形下,偏好稳定性与渐近稳定性等价。然而在更一般情况下,该等价性失效:我们构造了一个三玩家博弈,其偏好稳定集的跨度在动态上不稳定,表明偏好不足以作为动态稳定性的判断标准。随后,我们引入‘聚合偏离鲁棒性’这一基于收益的简单可验证条件,可保证任意纯策略组合跨度的渐近稳定性。
原文摘要 · Abstract (English)
We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players' learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames -- i.e. subsets of pure profiles obtained by restricting players' action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。