arXiv:2509.24239cs.LGcs.AI2025-09ACL被引 2

用国际象棋测试大模型真战略推理能力,发现多数表现不如人类业余选手。

ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models

  • 设计棋类对抗框架,让大模型在四种模式下对弈
  • 13个模型中无一击败人类水平的Maia-1100,部分输于随机走法
  • 微调后的Qwen3-8B表现显著提升,逼近顶尖推理模型

近期大型语言模型展现出强大的推理能力,但关键问题仍存:这些模型是否具备真正的战略推理能力,还是仅擅长模式识别?为解答此问题,我们提出ChessArena——一个基于国际象棋的评测基准。国际象棋要求战略思维、规则精准执行及复杂局面追踪。ChessArena是一个竞技框架,支持大模型在四种对战模式下相互对弈。我们在超过800局比赛中评估了13个大模型,涵盖基础理解、走法选择与棋题求解能力。结果揭示显著缺陷:无一模型能击败人类业余水平的Maia-1100,部分模型甚至输给了随机策略。我们还提出一个强基线:微调后的Qwen3-8B性能显著提升,接近更大规模的先进推理模型。

原文摘要 · Abstract (English)

Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess genuine strategic reasoning, or do they primarily excel at pattern recognition? To address this, we present ChessArena, a chess-based testbed for evaluating LLMs. Chess demands strategic reasoning, precise rule adherence, and the ability to track complex game states. ChessArena is a competitive framework where LLMs play against each other under four play modes. We evaluate 13 LLMs across over 800 games, testing basic understanding, move selection, and puzzle solving. Results reveal significant shortcomings: no model beats Maia-1100 (human amateur level), and some lose to random play. We also present a strong baseline: our fine-tuned Qwen3-8B substantially improves performance, approaching much larger state-of-the-art reasoning models.

战略推理大模型评测国际象棋

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。