用强化学习提升大模型下象棋的空间策略推理能力
Xiangqi-R1: Enhancing Spatial Strategic Reasoning in LLMs for Chinese Chess via Reinforcement Learning
- 基于五百万局棋谱+专家标注,分阶段训练70亿参数模型
- 移动合法性提升18%,分析准确率提高22%
- 适合研究通用战略智能与复杂棋类推理的学者
棋类游戏长期作为评估通用人工智能的基础基准。尽管大语言模型在通用推理方面表现卓越,但其在复杂、完全可观测棋类中至关重要的空间战略推理能力仍不足。本文以规则复杂、空间性强的中国象棋(Xiangqi)为挑战性测试平台,构建了一个包含五百万个棋局-走法对的大规模数据集,并加入专家标注与引擎评估。在此基础上,提出Xiangqi-R1,一个70亿参数的模型,采用多阶段训练框架。实验表明,通用大模型在该任务上表现不佳;而相比通用模型,Xiangqi-R1在移动合法性上提升18%,分析准确率提高22%。结果揭示了在复杂领域实现通用战略智能的可行路径。
原文摘要 · Abstract (English)
Game playing has long served as a fundamental benchmark for evaluating Artificial General Intelligence. While Large Language Models (LLMs) have demonstrated impressive capabilities in general reasoning, their effectiveness in spatial strategic reasoning, which is critical for complex and fully observable board games, remains insufficiently explored. In this work, we adopt Chinese Chess (Xiangqi) as a challenging and rich testbed due to its intricate rules and spatial complexity. To advance LLMs' strategic competence in such environments, we propose a training framework tailored to Xiangqi, built upon a large-scale dataset of five million board-move pairs enhanced with expert annotations and engine evaluations. Building on this foundation, we introduce Xiangqi-R1, a 7B-parameter model trained in multi-stage manner. Our Experimental results indicate that, despite their size and power, general-purpose LLMs struggle to achieve satisfactory performance in these tasks. Compared to general-purpose LLMs, Xiangqi-R1 greatly advances with an 18% rise in move legality and a 22% boost in analysis accuracy. Our results point to a promising path for creating general strategic intelligence in complex areas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。