arXiv:2506.10209cs.CLcs.AI2025-06EMNLP被引 4

用简单又新颖的井字棋游戏评测大模型的推理能力,发现顶尖模型在基础策略题上频频失分。

TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games

  • 设计四款人类易解的井字棋变体,通过程序化生成可验证的双人对战题
  • 主流大模型在该基准上平均比数学竞赛题低41%,长序列策略题表现更差
  • 适合评估模型真实战略推理能力,尤其关注长期规划与对手意图推断

大型推理模型(LRMs)在奥数级数学问题等任务中展现出卓越的推理能力,暗示其具备复杂推理潜力。然而,现有基准多集中于理工科领域,模型在更广泛任务中的推理表现仍缺乏探索。本文提出TTT-Bench,一个通过四款两人井字棋类游戏评估模型基本战略、空间与逻辑推理能力的新基准,这些游戏人类从幼年即可轻松解决。我们采用简单但可扩展的程序化方法生成可验证的双人博弈问题。尽管对人类而言极为简单,这些游戏需模型理解对手意图并分析棋盘空间布局才能取胜。我们评估了多种先进大模型,发现擅长高难度数学题的模型在这些简单游戏中频繁失败。进一步测试显示,所评估模型在TTT-Bench上的平均得分较MATH 500和AIME 2024分别下降41%和5%;更大模型虽以更短推理路径取得更高性能,但在涉及长期战略的简单新任务上仍普遍表现不佳。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiad-level mathematical problems, indicating evidence of their complex reasoning abilities. While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to reason correctly in broader task domains remains underexplored. In this work, we introduce \textbf{TTT-Bench}, a new benchmark that is designed to evaluate basic strategic, spatial, and logical reasoning abilities in LRMs through a suite of four two-player Tic-Tac-Toe-style games that humans can effortlessly solve from a young age. We propose a simple yet scalable programmatic approach for generating verifiable two-player game problems for TTT-Bench. Although these games are trivial for humans, they require reasoning about the intentions of the opponent, as well as the game board's spatial configurations, to ensure a win. We evaluate a diverse set of state-of-the-art LRMs, and \textbf{discover that the models that excel at hard math problems frequently fail at these simple reasoning games}. Further testing reveals that our evaluated reasoning models score on average $\downarrow$ 41\% \& $\downarrow$ 5\% lower on TTT-Bench compared to MATH 500 \& AIME 2024 respectively, with larger models achieving higher performance using shorter reasoning traces, where most of the models struggle on long-term strategic reasoning situations on simple and new TTT-Bench tasks.

推理评测策略游戏大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。