arXiv:2608.28884cs.AI2026-08

用游戏环境评测大模型建房能力,发现其真实工程缺陷。

MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

论文配图:MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft
图 1 · 摘自论文原文
  • 设计723条专家级自然语言指令,自动验证建造结果
  • 17类任务中大模型平均成功率不足60%,常出空间错误
  • 适合评估大模型在真实工程场景中的可靠性

我们提出MineCEraft(Minecraft建造工程基准),一个开源、易用的基准测试体系,用于系统评估大语言模型在Minecraft世界中执行建造任务的可靠性和局限性。该基准包含723条领域专家手工撰写的自然语言指令,覆盖17类不同任务,支持程序化可验证评估,提供安全可控的实验环境,用于检验大模型完成真实建造工程任务的能力。基于此基准,我们对当前最先进的大模型进行了深入评估,并开展了详细错误分析,揭示了大模型在应用于建造工程任务时的关键失败模式与实际挑战。

原文摘要 · Abstract (English)

We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.

大模型评测游戏智能自然语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。