arXiv:2411.17255cs.LGcs.AI2024-11被引 3

用大模型让机器人在Minecraft里自主设计并建造复杂结构。

APT: Architectural Planning and Text-to-Blueprint Construction Using Large Language Models for Open-World Agents

  • 通过思维链分解和多模态输入,让大模型生成可执行的建筑蓝图。
  • 能理解大量物品的位置与方向,成功构建带红石系统的复杂结构。
  • 记忆模块提升性能,还意外出现搭脚手架等类人规划行为。

我们提出APT框架,一种基于大语言模型(LLM)的先进方法,使自主代理能够在Minecraft环境中构建复杂且富有创意的结构。与以往专注于技能任务或依赖图像扩散模型生成体素结构的方法不同,本方法利用大模型固有的空间推理能力,结合思维链分解与多模态输入,生成详细建筑布局与蓝图,支持零样本或少样本场景下的执行。代理集成记忆与反思模块,实现持续学习、自适应优化与错误修正。为全面评估该领域表现,我们设计了一个包含多样化建造任务的基准测试,涵盖创造力、空间推理、规则遵守及多模态指令融合能力。实验使用多种GPT系列大模型后端与代理配置,结果表明代理能准确解析涉及大量物品及其位置、朝向的复杂指令,并成功构建具备红石系统等内部功能的结构。A/B测试显示,引入记忆模块显著提升性能,凸显其在积累经验与持续学习中的关键作用。此外,代理意外表现出搭脚手架等行为,揭示了未来大模型驱动代理通过子程序规划,自主发展类人问题解决策略的潜力。

原文摘要 · Abstract (English)

We present APT, an advanced Large Language Model (LLM)-driven framework that enables autonomous agents to construct complex and creative structures within the Minecraft environment. Unlike previous approaches that primarily concentrate on skill-based open-world tasks or rely on image-based diffusion models for generating voxel-based structures, our method leverages the intrinsic spatial reasoning capabilities of LLMs. By employing chain-of-thought decomposition along with multimodal inputs, the framework generates detailed architectural layouts and blueprints that the agent can execute under zero-shot or few-shot learning scenarios. Our agent incorporates both memory and reflection modules to facilitate lifelong learning, adaptive refinement, and error correction throughout the building process. To rigorously evaluate the agent's performance in this emerging research area, we introduce a comprehensive benchmark consisting of diverse construction tasks designed to test creativity, spatial reasoning, adherence to in-game rules, and the effective integration of multimodal instructions. Experimental results using various GPT-based LLM backends and agent configurations demonstrate the agent's capacity to accurately interpret extensive instructions involving numerous items, their positions, and orientations. The agent successfully produces complex structures complete with internal functionalities such as Redstone-powered systems. A/B testing indicates that the inclusion of a memory module leads to a significant increase in performance, emphasizing its role in enabling continuous learning and the reuse of accumulated experience. Additionally, the agent's unexpected emergence of scaffolding behavior highlights the potential of future LLM-driven agents to utilize subroutine planning and leverage the emergence ability of LLMs to autonomously develop human-like problem-solving techniques.

大模型Minecraft自主建造空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。