arXiv:2410.09701stat.MLcs.GT2024-10NeurIPS被引 2

预训练的Transformer模型可证明性地实现博弈策略学习,逼近纳什均衡。

Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models

  • 利用Transformer架构实现分布式与集中式博弈学习
  • 理论证明模型能在上下文内逼近零和博弈的纳什均衡
  • 适合对博弈论与大模型机制感兴趣的读者

基于Transformer架构的预训练模型在上下文学习(ICL)方面受到越来越多关注。尽管已有理论研究揭示了其在强化学习中的上下文学习能力,但以往成果主要集中于单智能体场景。本文进一步探索预训练Transformer模型在竞争性多智能体游戏中的上下文学习能力,即上下文博弈学习(ICGP)。针对经典的双人零和博弈,本文提供了理论保证,证明预训练Transformer可在上下文方式下,分别在分布式与集中式学习设置中,逼近纳什均衡。作为证明的关键部分,构建性结果表明:Transformer架构足够丰富,能够实现经典多智能体博弈算法,特别是分布式V-learning和集中式VI-ULCB。

原文摘要 · Abstract (English)

The in-context learning (ICL) capability of pre-trained models based on the transformer architecture has received growing interest in recent years. While theoretical understanding has been obtained for ICL in reinforcement learning (RL), the previous results are largely confined to the single-agent setting. This work proposes to further explore the in-context learning capabilities of pre-trained transformer models in competitive multi-agent games, i.e., in-context game-playing (ICGP). Focusing on the classical two-player zero-sum games, theoretical guarantees are provided to demonstrate that pre-trained transformers can provably learn to approximate Nash equilibrium in an in-context manner for both decentralized and centralized learning settings. As a key part of the proof, constructional results are established to demonstrate that the transformer architecture is sufficiently rich to realize celebrated multi-agent game-playing algorithms, in particular, decentralized V-learning and centralized VI-ULCB.

博弈学习Transformer纳什均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。