构建Minecraft多模态多智能体协作评测基准,测试智能体泛化能力。
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
- 基于Minecraft构建5.5万种多模态任务变体,支持视觉、语言等多源输入
- 提供程序生成的专家演示数据,用于模仿学习与泛化评估
- 揭示当前模型在新目标、新场景、多智能体数量下仍存显著局限
协作是社会的核心。在现实世界中,人类队友利用多感官数据应对动态环境中的复杂任务。在视觉丰富且互动频繁的环境中,具身智能体需理解多模态观测与任务指令。为评估通用多模态协作智能体的表现,我们提出TeamCraft——一个基于开放世界游戏Minecraft构建的多模态多智能体基准。该基准包含55,000个由多模态提示定义的任务变体,通过程序生成的专家示范数据支持模仿学习,并设计了严格评估协议以检验模型泛化能力。我们进行了广泛分析,揭示现有方法在应对新目标、新场景及未见智能体数量时仍面临严峻挑战。这些发现凸显了该领域亟待进一步研究。TeamCraft平台与数据集已公开:https://github.com/teamcraft-bench/teamcraft。
原文摘要 · Abstract (English)
Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essential for embodied agents collaborating in visually-rich environments replete with dynamic interactions to understand multi-modal observations and task specifications. To evaluate the performance of generalizable multi-modal collaborative agents, we present TeamCraft, a multi-modal multi-agent benchmark built on top of the open-world video game Minecraft. The benchmark features 55,000 task variants specified by multi-modal prompts, procedurally-generated expert demonstrations for imitation learning, and carefully designed protocols to evaluate model generalization capabilities. We also perform extensive analyses to better understand the limitations and strengths of existing approaches. Our results indicate that existing models continue to face significant challenges in generalizing to novel goals, scenes, and unseen numbers of agents. These findings underscore the need for further research in this area. The TeamCraft platform and dataset are publicly available at https://github.com/teamcraft-bench/teamcraft.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。