arXiv:2505.13638cs.LGcs.CL2025-05

构建复杂棋类环境4Hammer,用于长时序强化学习与大模型评估

4Hammer: a board-game reinforcement learning environment for the hour long time frame

  • 基于战锤40K设计数字孪生棋类环境,支持长周期策略博弈
  • 模拟超过50页自然语言规则,考验模型对复杂状态的持续理解能力
  • 适合评估大模型在超长任务中的规划与推理能力

大型语言模型在短时任务中表现优异,但在需长时间持续决策的任务中表现不佳。尽管已有涵盖长时任务的数据集(如软件工程或电子游戏),但针对强化学习和大模型评估而设计的复杂棋类游戏实现仍十分稀少。为填补这一空白,我们提出4Hammer强化学习环境,这是一个战锤40,000(Warhammer 40,000)部分规则的数字孪生仿真系统。战锤40,000规则极为复杂,要求玩家阅读并理解超过50页的详细自然语言规则,掌握己方与对手单位间的交互关系,并独立追踪和沟通不断演变的游戏状态。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated strong performance on tasks with short time frames, but struggle with tasks requiring longer durations. While datasets covering extended-duration tasks, such as software engineering tasks or video games, do exist, there are currently few implementations of complex board games specifically designed for reinforcement learning and LLM evaluation. To address this gap, we propose the 4Hammer reinforcement learning environment, a digital twin simulation of a subset of Warhammer 40,000-a complex, zero-sum board game. Warhammer 40,000 features intricate rules, requiring human players to thoroughly read and understand over 50 pages of detailed natural language rules, grasp the interactions between their game pieces and those of their opponents, and independently track and communicate the evolving game state.

强化学习长时序决策棋类游戏大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。