大模型在德州扑克中自发学习到环境的随机状态表示。
Emergent World Beliefs: Exploring Transformers in Stochastic Games
- 用扑克历史数据预训练GPT模型,探测其内部激活
- 模型自发学习手牌等级和胜率等确定与随机特征
- 无需显式指令,就能构建对不确定环境的信念表征
基于Transformer的大语言模型(LLMs)已在编程、棋类等复杂任务中展现出强大的推理能力。已有研究证明,这类模型在完全信息博弈中能发展出对环境状态的隐式表征。本文将这一研究拓展至不完全信息领域,以德州扑克作为部分可观测马尔可夫决策过程(POMDP)的典型范例。我们在扑克手牌历史(PHH)数据上对一个GPT风格模型进行预训练,并分析其内部激活。结果表明,模型在无显式指导的情况下,既学习到了确定性结构(如手牌等级),也捕捉到随机性特征(如胜率)。通过主要使用非线性探测器,我们验证了这些表征可解码且与理论上的信念状态高度相关,说明LLMs能够自发构建对德州扑克这类随机环境的认知模型。
原文摘要 · Abstract (English)
Transformer-based large language models (LLMs) have demonstrated strong reasoning abilities across diverse fields, from solving programming challenges to competing in strategy-intensive games such as chess. Prior work has shown that LLMs can develop emergent world models in games of perfect information, where internal representations correspond to latent states of the environment. In this paper, we extend this line of investigation to domains of incomplete information, focusing on poker as a canonical partially observable Markov decision process (POMDP). We pretrain a GPT-style model on Poker Hand History (PHH) data and probe its internal activations. Our results demonstrate that the model learns both deterministic structure, such as hand ranks, and stochastic features, such as equity, without explicit instruction. Furthermore, by using primarily nonlinear probes, we demonstrated that these representations are decodeable and correlate with theoretical belief states, suggesting that LLMs are learning their own representation of the stochastic environment of Texas Hold'em Poker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。