arXiv:2503.13980cs.LG2025-03被引 3

用合成数据训练大模型下棋,提升复杂决策能力

Empowering LLMs in Decision Games through Algorithmic Data Synthesis

  • 通过合成斗地主和围棋数据,定制化后训练大模型
  • 新模型在各自游戏中达到可竞争水平,胜率显著提升
  • 成果适合想提升推理能力的AI研究者和开发者

大型语言模型(LLMs)在众多领域展现出强大能力,但在复杂推理与决策任务中仍表现不足。决策游戏因其多维度推理需求,成为评估和增强LLM推理能力的理想场景。本文探索了通过针对性后训练使LLMs掌握复杂决策游戏的可能性。为此,设计数据合成策略,从经典游戏斗地主(Doudizhu)和围棋(Go)中构建大规模离线数据集,并开发一系列技术将这些数据有效融入LLM训练,最终形成两个新智能体:Mastermind-Dou 和 Mastermind-Go。实验表明,这些基于大模型的智能体在对应游戏中表现出竞争力。此外,研究还发现引入决策数据能提升模型部分推理能力,为优化大模型数据收集与合成策略提供了重要启示。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have exhibited impressive capabilities across numerous domains, yet they often struggle with complex reasoning and decision-making tasks. Decision-making games, which inherently require multifaceted reasoning logic, serve as ideal sandboxes for evaluating and enhancing the reasoning abilities of LLMs. In this work, we first explore whether LLMs can master complex decision-making games through targeted post-training. To this end, we design data synthesis strategies and curate extensive offline datasets from two classic games, Doudizhu and Go. We further develop a suite of techniques to effectively incorporate this data into LLM training, resulting in two novel agents: Mastermind-Dou and Mastermind-Go. Our experimental results demonstrate that these Mastermind LLMs achieve competitive performance in their respective games. Additionally, we explore whether integrating decision-making data can enhance the general reasoning abilities of LLMs. Our findings suggest that such post-training improves certain aspects of reasoning, providing valuable insights for optimizing LLM data collection and synthesis strategies.

决策推理数据合成大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。