提升Minecraft对话中指令跟随能力,解决数据少与空间推理弱问题。
BAP v2: An Enhanced Task Framework for Instruction Following in Minecraft Dialogues
- 构建新评估框架,优化数据与指标以更真实反映模型表现。
- 合成多种模拟数据训练,显著提升模型在简单任务上的表现。
- 提出Llama-CRAFTS模型,F1达53.0,仍暴露空间推理短板。
构建可理解语言、感知环境并在物理世界中行动的交互智能体是人工智能长期目标。Minecraft协作建造任务(MCBT)提供了一个双人游戏平台,其中建筑师(A)指导建造者(B)在3D方块世界中搭建目标结构,是实现该目标的理想场景。本文聚焦建造者动作预测(BAP)子任务:在多模态游戏中预测建造者的行动,这是一个充满挑战的具身指令跟随测试平台,且训练数据有限。我们全面重新审视该任务,提出BAP v2以解决评估、训练数据和建模三大难题。具体而言,我们设计了更清晰的测试集和更公平、更具洞察力的评估指标,揭示空间推理是主要性能瓶颈。为缓解数据稀缺并培养模型基本空间能力,我们生成了多种合成的MCBT数据。发现当前基于大模型的最先进模型在人类对话数据上表现良好,但在合成数据上失败;但通过合成数据训练后,模型在各项任务上均有提升。我们还提出新模型Llama-CRAFTS,利用更丰富的输入表示,在BAP v2任务上取得53.0的F1分数,并在合成数据上表现优异。尽管相比之前工作有6分提升,但仍凸显任务难度,确立了BAP v2作为未来研究的肥沃土壤,也为评估当前纯文本大模型在具身任务中的空间能力提供了有效基准。
原文摘要 · Abstract (English)
Developing interactive agents that can understand language, perceive their surroundings, and act within the physical world is a long-standing goal of AI research. The Minecraft Collaborative Building Task (MCBT) (Narayan-Chen, Jayannavar, and Hockenmaier 2019), a two-player game in which an Architect (A) instructs a Builder (B) to construct a target structure in a simulated 3D Blocks World environment, offers a rich platform to work towards this goal. In this work, we focus on the Builder Action Prediction (BAP) subtask: predicting B's actions in a multimodal game context (Jayannavar, Narayan-Chen, and Hockenmaier 2020) - a challenging testbed for grounded instruction following, with limited training data. We holistically re-examine this task and introduce BAP v2 to address key challenges in evaluation, training data, and modeling. Specifically, we define an enhanced evaluation benchmark, featuring a cleaner test set and fairer, more insightful metrics that also reveal spatial reasoning as the primary performance bottleneck. To address data scarcity and to teach models basic spatial skills, we generate different types of synthetic MCBT data. We observe that current, LLM-based SOTA models trained on the human BAP dialogues fail on these simpler, synthetic BAP ones, but show that training models on this synthetic data improves their performance across the board. We also introduce a new SOTA model, Llama-CRAFTS, which leverages richer input representations, and achieves an F1 score of 53.0 on the BAP v2 task and strong performance on the synthetic data. While this result marks a notable 6 points improvement over previous work, it also underscores the task's remaining difficulty, establishing BAP v2 as a fertile ground for future research, and providing a useful measure of the spatial capabilities of current text-only LLMs in such embodied tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。