构建首个系统性评估大模型下棋理解能力的基准测试
ChessQA: Evaluating Large Language Models for Chess Understanding
- 设计五类任务覆盖从规则到高阶概念的棋艺抽象层次
- 多模型测试发现所有类别均存在持续性弱点
- 支持动态更新,适合研究模型进化与对比
国际象棋是评估大语言模型(LLM)推理、建模与抽象能力的理想测试平台,因其结构清晰、目标明确且技能水平跨度大。然而,现有评估方法零散且范围狭窄,难以准确衡量模型在不同规模、微调方法或架构下的棋艺理解能力。本文提出ChessQA,一个涵盖五类任务(结构理解、战术模式、短程战术、局面判断、语义描述)的综合性评估基准,对应棋手从掌握基础规则到理解高阶概念的渐进认知过程。该基准不仅全面超越以往仅评价走法质量的局限,还提供可控、一致的诊断与比较环境。此外,ChessQA具备动态性,可随模型进步持续更新提示、答案与数据构造脚本。对多种主流LLM的评估显示,各任务类别均存在持久性缺陷,并按类别提供了结果与错误分析。代码、定期更新的数据集及公开排行榜将向社区开放,以支持后续研究。
原文摘要 · Abstract (English)
Chess provides an ideal testbed for evaluating the reasoning, modeling, and abstraction capabilities of large language models (LLMs), as it has well-defined structure and objective ground truth while admitting a wide spectrum of skill levels. However, existing evaluations of LLM ability in chess are ad hoc and narrow in scope, making it difficult to accurately measure LLM chess understanding and how it varies with scale, post-training methodologies, or architecture choices. We present ChessQA, a comprehensive benchmark that assesses LLM chess understanding across five task categories (Structural, Motifs, Short Tactics, Position Judgment, and Semantic), which approximately correspond to the ascending abstractions that players master as they accumulate chess knowledge, from understanding basic rules and learning tactical motifs to correctly calculating tactics, evaluating positions, and semantically describing high-level concepts. In this way, ChessQA captures a more comprehensive picture of chess ability and understanding, going significantly beyond the simple move quality evaluations done previously, and offers a controlled, consistent setting for diagnosis and comparison. Furthermore, ChessQA is inherently dynamic, with prompts, answer keys, and construction scripts that can evolve as models improve. Evaluating a range of contemporary LLMs, we find persistent weaknesses across all five categories and provide results and error analyses by category. We will release the code, periodically refreshed datasets, and a public leaderboard to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。