arXiv:2606.15684cs.AI2026-06被引 1

用Minecraft搭建多智能体协作测试平台,验证实时协同的挑战。

Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft

论文配图:Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft
图 1 · 摘自论文原文
  • 构建可动态生成任务的框架,支持声明式配置
  • 大模型在部分观测下协作失败率高,远低于全局知识基准
  • 适合研究实时多智能体系统与鲁棒协同算法的学者

我们提出TickingCollabBench,一个基于Minecraft的多智能体时间敏感互补协作基准。该基准体现现实协作的四大特征:智能体异质性、强制协作、动态环境和严格的实时约束及失败风险。为此,我们开发了TickingCollab框架,支持生成多样化动态环境,并抽象Minecraft原始API,实现以声明式YAML格式定义任务事件。在此基础上,设计了一个可行性感知的自动化基准生成流水线:大语言模型(LLM)生成结构多样任务配置,可行性验证器利用近似约束过滤无效配置。评估显示,语言延迟和在部分可观测性与智能体异质性下的协调难度,导致大模型在动态环境中频繁失败,显著落后于全局知识代理。

原文摘要 · Abstract (English)

We present TickingCollabBench, a Minecraft-based multi-agent benchmark for a novel class of time-sensitive complementary collaboration tasks. Our benchmark reflects four core characteristics of real-world collaboration: agent heterogeneity, mandatory collaboration, dynamic environments, and strict real-time constraints with failure risks. To enable this, we develop the TickingCollab framework, which supports the generation of diverse dynamic environments and abstracts Minecraft's primitive APIs to enable declarative YAML task specifications for composing these events. Building on this, we design a feasibility-aware automated benchmark generation pipeline, where an LLM drafts structurally diverse task configurations and feasibility verifier filters out invalid ones using approximate constraints. Evaluations demonstrate that lang latency and inherent difficulty of coordinating under partial observability and agent heterogeneity cause LLMs to frequently fail under dynamic environments and fall significantly short of a global-knowledge oracle.

多智能体实时协作游戏基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。