用国际象棋游戏验证大模型能否自发构建世界模型
Revisiting the Othello World Model Hypothesis
- 用棋局序列训练7个大模型预测下一步
- 所有模型达99%准确率,且学到相似棋盘特征
- 为大模型具身认知能力提供更强证据
Li等人(2023)以国际象棋为案例,检验GPT-2构建世界模型的能力,后续被Nanda等人(2023b)跟进。本文简要回顾原实验,并扩展至更多语言模型与更全面的探针分析。具体而言,我们分析国际象棋棋局序列,训练模型根据先前步数预测下一步。在七个语言模型(GPT-2、T5、Bart、Flan-T5、Mistral、LLaMA-2和Qwen2.5)上评估国际象棋任务,发现这些模型不仅学会下国际象棋,还自发构建了棋盘布局表征。所有模型在无监督对齐任务中达到最高99%准确率,且学习到的棋盘特征高度相似。这比以往研究提供了更有力的证据支持国际象棋世界模型假说。
原文摘要 · Abstract (English)
Li et al. (2023) used the Othello board game as a test case for the ability of GPT-2 to induce world models, and were followed up by Nanda et al. (2023b). We briefly discuss the original experiments, expanding them to include more language models with more comprehensive probing. Specifically, we analyze sequences of Othello board states and train the model to predict the next move based on previous moves. We evaluate seven language models (GPT-2, T5, Bart, Flan-T5, Mistral, LLaMA-2, and Qwen2.5) on the Othello task and conclude that these models not only learn to play Othello, but also induce the Othello board layout. We find that all models achieve up to 99% accuracy in unsupervised grounding and exhibit high similarity in the board features they learned. This provides considerably stronger evidence for the Othello World Model Hypothesis than previous works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。