用游戏规则测试大模型逻辑推理,发现越往后越容易出错。
Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing
- 在规则明确的游戏环境中评估大模型的推理能力
- 多数模型在多步推理中表现良好,但步数增多时性能下降
- 揭示了模型常犯的规则幻觉、冗余状态等错误
本文从新视角考察大型语言模型(LLMs)的推理能力,聚焦其在形式化规则环境中的表现。我们在一系列前向模拟任务(包括下一步/多步状态推演与合法动作生成)上,对四款模型(Gemini 2.5 Pro 及 Flash 版本、Llama 3.3 70B、GPT-OSS 120B)进行了评估,涵盖通过通用游戏玩法(GGP)实例呈现的多样化推理问题。除了报告个体实例的表现外,我们基于40个结构特征对游戏进行分类,并分析这些特征与模型性能之间的相关性。此外,我们研究了不同游戏混淆方式的影响,以评估语言语义在游戏定义中的作用,以及模型在训练中可能接触特定游戏的影响。主要结果表明,三款模型在多数实验设置中表现良好,但随着评估视野扩大(即游戏步数增加),性能出现下降。对模型表现的案例分析揭示了在逻辑推理任务中常见的错误模式,如虚构规则、冗余状态事实或语法错误。总体而言,本文报告了当代模型在形式化推理能力上的显著进展。
原文摘要 · Abstract (English)
This paper examines the reasoning capabilities of Large Language Models (LLMs) from a novel perspective, focusing on their ability to operate within formally specified, rule-governed environments. We evaluate four LLMs (Gemini 2.5 Pro and Flash variants, Llama 3.3 70B and GPT-OSS 120B) on a suite of forward-simulation tasks-including next / multistep state formulation, and legal action generation-across a diverse set of reasoning problems illustrated through General Game Playing (GGP) game instances. Beyond reporting instance-level performance, we characterize games based on 40 structural features and analyze correlations between these features and LLM performance. Furthermore, we investigate the effects of various game obfuscations to assess the role of linguistic semantics in game definitions and the impact of potential prior exposure of LLMs to specific games during training. The main results indicate that three of the evaluated models generally perform well across most experimental settings, with performance degradation observed as the evaluation horizon increases (i.e., with a higher number of game steps). Detailed case-based analysis of the LLM performance provides novel insights into common reasoning errors in the considered logic-based problem formulation, including hallucinated rules, redundant state facts, or syntactic errors. Overall, the paper reports clear progress in formal reasoning capabilities of contemporary models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。