构建首个覆盖144种博弈类型的系统化评测基准,测试大模型的战略推理能力。
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
- 基于2×2博弈拓扑构建144类经典博弈,结合故事化场景提升多样性
- 在序列、并行、嵌套结构中验证大模型推理一致性,准确率不足60%
- 专为高阶模型设计可扩展框架,适合评估思维链与心智理论能力
大语言模型在推理任务中的应用快速发展,战略推理日益受到关注。现有研究多集中于少数博弈类型,覆盖范围有限,且经典场景存在数据泄露风险,基准缺乏可扩展性。为此,我们提出TMGBench,涵盖罗宾逊-戈弗斯拓扑中所有144种2×2博弈类型,并构造经典博弈;同时为每类博弈生成多样化高质量的故事化场景。为支持持续评估,将这些博弈作为原子单元,通过序列、并行、嵌套结构组合成更复杂形式。对主流大模型的全面评估显示,其战略推理准确性和一致性仍不足,理论心智能力差异显著。SOTA模型如o3-mini、Qwen3和deepseek-reasoner在复杂结构中表现受限,凸显了该基准的挑战性。
原文摘要 · Abstract (English)
The rapid advancement of large language models has accelerated their application in reasoning, with strategic reasoning drawing increasing attention. To evaluate the strategic reasoning capabilities of LLMs, game theory, with its concise structure, has become the preferred approach for many researchers. However, current research typically focuses on a limited selection of games, resulting in low coverage of game types. Additionally, classic game scenarios carry risks of data leakage, and the benchmarks used often lack extensibility, rendering them inadequate for evaluating state-of-the-art models. To address these challenges, we propose TMGBench, characterized by comprehensive game type coverage, diverse scenarios and flexible game organization. Specifically, we incorporate all 144 game types summarized by the Robinson-Goforth topology of 2x2 games, constructed as classic games in our benchmark; we also synthetize diverse, higher-quality game scenarios for each classic game, which we refer to as story-based games. Lastly, to provide a sustainable evaluation framework adaptable to increasingly powerful LLMs, we treat the aforementioned games as atomic units and organize them into more complex forms through sequential, parallel, and nested structures. We conducted a comprehensive evaluation of mainstream LLMs, covering tests on rational reasoning, reasoning robustness, Theory-of-Mind capabilities, and reasoning in complex game forms. The results revealed LLMs still have flaws in the accuracy and consistency of strategic reasoning processes, and their levels of mastery over Theory-of-Mind also vary. Additionally, SOTA models like o3-mini, Qwen3 and deepseek-reasoner, were also evaluated across the sequential, parallel, and nested game structures while the results highlighted the challenges posed by TMGBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。