arXiv:2502.20770cs.GTcs.LG2025-02ICML被引 1

通过游戏反复交互,让优化器引导无悔学习者达到均衡。

Learning to Steer Learners in Games

  • 优化器通过反向推断学习者的收益结构来引导其行为。
  • 若学习者使用梯度上升或已知正则的随机镜面上升,可有效引导至均衡。
  • 适用于了解学习算法细节但不知收益函数的博弈场景。

我们研究在重复双人有限动作博弈中,通过反复交互来利用学习算法的问题。具体而言,一个优化器试图在不知道学习者收益函数的情况下,将其引导至斯塔克尔伯格均衡。我们首先证明,如果优化器只知道学习者使用的是广义无悔算法,则该目标无法实现。这表明优化器需要更多关于学习者目标或算法的信息才能成功引导。基于此直觉,我们将问题转化为优化器恢复学习者收益结构的过程。若学习者使用的是更小类别的算法,我们通过两个例子验证了该方法的有效性:一是学习者采用梯度上升算法,二是使用已知正则化项和步长的随机镜面上升算法。

原文摘要 · Abstract (English)

We consider the problem of learning to exploit learning algorithms through repeated interactions in games. Specifically, we focus on the case of repeated two player, finite-action games, in which an optimizer aims to steer a no-regret learner to a Stackelberg equilibrium without knowledge of its payoffs. We first show that this is impossible if the optimizer only knows that the learner is using an algorithm from the general class of no-regret algorithms. This suggests that the optimizer requires more information about the learner's objectives or algorithm to successfully exploit them. Building on this intuition, we reduce the problem for the optimizer to that of recovering the learner's payoff structure. We demonstrate the effectiveness of this approach if the learner's algorithm is drawn from a smaller class by analyzing two examples: one where the learner uses an ascent algorithm, and another where the learner uses stochastic mirror ascent with known regularizer and step sizes.

博弈学习无悔算法引导策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。