用专家标注的棋局策略与战术训练大模型,提升其象棋推理能力
Explore the Reasoning Capability of LLMs in the Chess Testbed
- 引入专家标注的策略与战术,增强模型棋局理解
- 在100万棋位数据上微调后,表现优于GPT、Claude等商用模型
- 语言解释能有效提升模型长时复杂推理能力,适合棋类算法研究者
推理是人类智能的核心能力。近年来,随着大规模数据集的出现,预训练的大规模语言模型展现出新的推理能力。然而,这些模型在长期、复杂的推理任务(如下棋)中仍表现不佳。基于专家棋手同时运用长期战略与短期战术并辅以语言解释的双轨模式,我们提出通过整合标注的策略与战术来提升大语言模型在象棋中的推理能力。具体地,我们构建了名为MATE的数据集,包含100万个棋局位置,每个位置均有专家标注的策略与战术候选步。我们在LLaMA-3-8B模型上进行微调,并在选择更优走法的任务中与当前领先的商用语言模型(GPT、Claude、Gemini)进行对比。实验表明,我们的模型表现优于上述所有模型。结果还发现,语言解释能够显著增强大语言模型的推理能力。
原文摘要 · Abstract (English)
Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing chess. Based on the observation that expert chess players employ a dual approach combining long-term strategic play with short-term tactical play along with language explanation, we propose improving the reasoning capability of large language models in chess by integrating annotated strategy and tactic. Specifically, we collect a dataset named MATE, which consists of 1 million chess positions with candidate moves annotated by chess experts for strategy and tactics. We finetune the LLaMA-3-8B model and compare it against state-of-the-art commercial language models in the task of selecting better chess moves. Our experiments show that our models perform better than GPT, Claude, and Gemini models. We find that language explanations can enhance the reasoning capability of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。