arXiv:2503.09617cs.MAcs.CL2025-03NeurIPS被引 5

用游戏《工厂突袭》测试大模型长期规划与资源优化能力

Factorio Learning Environment

  • 基于游戏构建可扩展的开放环境,考验智能体长期规划与程序合成能力
  • 大模型在短程任务中表现尚可,但无法在资源受限环境下有效纠错
  • 适合研究长时序推理、自动化系统设计的学者和开发者参考

大型语言模型正快速饱和现有基准,亟需新的开放式评估。我们提出基于游戏《工厂突袭》的因子学习环境(FLE),用于测试智能体在长期规划、程序合成和资源优化方面的能力。该环境提供指数级增长的挑战——从基础自动化到每秒处理数百万资源单位的复杂工厂。我们设定两种场景:(1) 实验室模式,包含八个结构化任务与固定资源;(2) 开放模式,在程序生成的地图上无限制地建造最大工厂。在两种设置下,我们发现模型仍缺乏强空间推理能力。在实验室模式中,大模型展现出一定的短期操作能力,但在资源受限环境中难以有效运行,反映其错误分析能力不足;在开放模式中,尽管模型能发现提升增长的自动化策略(如电力钻探),却无法实现复杂自动化(如电子电路制造)。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are rapidly saturating existing benchmarks, necessitating new open-ended evaluations. We introduce the Factorio Learning Environment (FLE), based on the game of Factorio, that tests agents in long-term planning, program synthesis, and resource optimization. FLE provides exponentially scaling challenges -- from basic automation to complex factories processing millions of resource units per second. We provide two settings: (1) lab-play consisting of eight structured tasks with fixed resources, and (2) open-play with the unbounded task of building the largest factory on an procedurally generated map. We demonstrate across both settings that models still lack strong spatial reasoning. In lab-play, we find that LLMs exhibit promising short-horizon skills, yet are unable to operate effectively in constrained environments, reflecting limitations in error analysis. In open-play, while LLMs discover automation strategies that improve growth (e.g electric-powered drilling), they fail to achieve complex automation (e.g electronic-circuit manufacturing).

强化学习长期规划游戏环境大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。