arXiv:2503.21683cs.AIcs.CL2025-03

用大模型模拟人类下五子棋,自对弈强化学习提升决策能力

LLM-Gomoku: A Large Language Model-Based System for Strategic Gomoku with Self-Play and Reinforcement Learning

  • 通过读板、懂规则、选策略、评位置实现类人决策
  • 自对弈训练后棋力显著提升,非法落子率降低,评估速度加快
  • 适合研究大模型博弈与强化学习融合的开发者参考

近年来,大语言模型(LLMs)在自然语言处理领域取得显著进展,展现出强大的生成、理解与推理能力,已应用于教育、智能决策和游戏等领域。然而,如何有效利用大模型进行五子棋这类策略性游戏的规划与决策仍是挑战。本研究旨在构建一个基于大语言模型的五子棋人工智能系统,模拟人类学棋过程。系统设计包含‘读板’、‘懂规则’、‘选策略’、‘评位置’四大能力,并通过自对弈与强化学习持续优化。实验结果表明,该方法显著提升了落子位置选择精度,解决了生成非法落子的问题,并通过并行位置评估减少了处理时间。经过大量自对弈训练,模型的五子棋对弈能力得到明显增强。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have shown significant advancements in natural language processing (NLP), with strong capa-bilities in generation, comprehension, and rea-soning. These models have found applications in education, intelligent decision-making, and gaming. However, effectively utilizing LLMs for strategic planning and decision-making in the game of Gomoku remains a challenge. This study aims to develop a Gomoku AI system based on LLMs, simulating the human learning process of playing chess. The system is de-signed to understand and apply Gomoku strat-egies and logic to make rational decisions. The research methods include enabling the model to "read the board," "understand the rules," "select strategies," and "evaluate positions," while en-hancing its abilities through self-play and rein-forcement learning. The results demonstrate that this approach significantly improves the se-lection of move positions, resolves the issue of generating illegal positions, and reduces pro-cess time through parallel position evaluation. After extensive self-play training, the model's Gomoku-playing capabilities have been notably enhanced.

大模型五子棋AI自对弈强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。