构建Minecraft场景下的规划评估数据集,测试大模型的决策与推理能力。
Plancraft: an evaluation dataset for planning with LLM agents
- 基于Minecraft界面设计多模态评测任务,支持文本与视觉输入。
- 包含可解与不可解任务,检验模型判断问题可行性能力。
- 对比开源与闭源模型表现,揭示当前大模型在复杂规划中的不足。
我们提出Plancraft,一个用于大语言模型智能体规划能力的多模态评估数据集。该数据集基于Minecraft Crafting GUI构建,提供纯文本与多模态两种接口。引入Minecraft Wiki以评估工具使用和检索增强生成(RAG)能力,并设计人工规划器与Oracle Retriever,用于剖析现代智能体架构中各组件的作用。为评估决策能力,数据集还包含一组故意设计为无解的任务,模拟真实场景中需判断任务是否可行的挑战。我们对开源与闭源大模型进行了基准测试,并与人工规划器进行性能和效率对比。结果表明,当前大模型与视觉语言模型在Plancraft提出的规划任务中仍表现不佳,本文据此提出改进建议。
原文摘要 · Abstract (English)
We present Plancraft, a multi-modal evaluation dataset for LLM agents. Plancraft has both a text-only and multi-modal interface, based on the Minecraft crafting GUI. We include the Minecraft Wiki to evaluate tool use and Retrieval Augmented Generation (RAG), as well as a handcrafted planner and Oracle Retriever, to ablate the different components of a modern agent architecture. To evaluate decision-making, Plancraft also includes a subset of examples that are intentionally unsolvable, providing a realistic challenge that requires the agent not only to complete tasks but also to decide whether they are solvable at all. We benchmark both open-source and closed-source LLMs and compare their performance and efficiency to a handcrafted planner. Overall, we find that LLMs and VLMs struggle with the planning problems that Plancraft introduces, and offer suggestions on how to improve their capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。