arXiv:2602.15867cs.CLcs.AI2026-02

测试顶尖大模型在经典文字冒险游戏中的表现,发现其解题能力严重不足。

Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?

  • 用1977年经典文字游戏Zork评估大模型的推理与规划能力。
  • 最优秀模型仅得75分(满分350),平均完成度不足10%。
  • 即使给详细指令或开启深度思考,表现也无提升,暴露认知短板。

本文通过评估当代大型语言模型(LLMs)在1977年发布的经典文字冒险游戏Zork中的表现,检验其问题解决与推理能力。游戏以对话式结构为特点,可作为评估基于LLM聊天机器人理解自然语言描述并生成恰当动作序列的可控环境。我们测试了主流专有模型ChatGPT、Claude和Gemini在最少与详细指令下的表现,以得分作为主要评估指标。结果显示,所有模型平均完成度不足10%,最佳模型Claude Opus 4.5仅获得约75分(满分350)。值得注意的是,提供详细指令或启用“扩展思考”功能均未带来性能提升。定性分析揭示模型存在根本缺陷:重复失败动作表明缺乏自我反思能力,策略执行不一致,且虽能访问历史对话仍无法从过往尝试中学习。这些发现表明当前大模型在文本类游戏领域存在显著的认知与规划局限,对其推理本质提出质疑。

原文摘要 · Abstract (English)

In this positioning paper, we evaluate the problem-solving and reasoning capabilities of contemporary Large Language Models (LLMs) through their performance in Zork, the seminal text-based adventure game first released in 1977. The game's dialogue-based structure provides a controlled environment for assessing how LLM-based chatbots interpret natural language descriptions and generate appropriate action sequences to succeed in the game. We test the performance of leading proprietary models - ChatGPT, Claude, and Gemini - under both minimal and detailed instructions, measuring game progress through achieved scores as the primary metric. Our results reveal that all tested models achieve less than 10% completion on average, with even the best-performing model (Claude Opus 4.5) reaching only approximately 75 out of 350 possible points. Notably, providing detailed game instructions offers no improvement, nor does enabling ''extended thinking''. Qualitative analysis of the models' reasoning processes reveals fundamental limitations: repeated unsuccessful actions suggesting an inability to reflect on one's own thinking, inconsistent persistence of strategies, and failure to learn from previous attempts despite access to conversation history. These findings suggest substantial limitations in current LLMs' metacognitive abilities and problem-solving capabilities within the domain of text-based games, raising questions about the nature and extent of their reasoning capabilities.

大模型评测文本游戏推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。