提出多智能体协作任务评估基准,揭示协同能力因任务类型而异。
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning

- 基于897个验证实例,按四大协作模式设计评估体系
- 发现模型整体表现好但各协作能力不均衡
- 适合研究多智能体系统协同机制的学者使用
由多模态大语言模型驱动的智能体系统近年来发展迅速,但现有具身智能体基准仍缺乏对多智能体协作的细粒度诊断。多数基准或聚焦单智能体任务完成,或仅以总体任务成功率总结多智能体行为,难以揭示重复工作、顺序违规、资源争用和交接不同步等协作失败问题。本文提出CoCoBench,一个针对可执行家庭任务中多智能体具身协作的构造级评估基准。该基准包含897个经过奥拉验证的实例,围绕四大常见协作构造:任务分配、顺序排列、互斥约束和交接协调。除任务成功率外,还提供构造级评分以衡量协作有效性。我们在11个主流MLLMs上评估了不同协作模式、观测输入和智能体数量下的表现。结果显示,协作能力高度依赖具体构造:整体表现优异并不意味着各类协作能力均衡。这些发现为设计针对性模型架构和提升多智能体协作能力指明了新方向。
原文摘要 · Abstract (English)
Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination. Most benchmarks either focus on single-agent task completion or summarize multi-agent behavior with overall task success rates, which can obscure coordination failures such as duplicated work, violations of ordering constraints, resource contention, and desynchronized handoffs. In this paper, we introduce CoCoBench, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks. CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively. We evaluate 11 leading MLLMs across different coordination modes, observation inputs, and numbers of agents. The results show that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types. These findings point to new directions for designing targeted model architectures and improving multi-agent coordination ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。