评测多模态智能体在真实环境中的协作能力,发现沟通是关键,但协作效果取决于团队规模与模型能力。
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

- 构建跨任务、双结构、三模式的多智能体协作评测平台
- 协作提升任务完成率,但需权衡收益与协调复杂度
- 通信对协作增益至关重要,且最优模式随团队大小变化
近期多模态大语言模型(MLLMs)在具身智能体中展现出巨大潜力,但其在视觉引导环境中的协作能力仍待深入探索。为填补这一空白,我们提出MECoBench,一个涵盖多样化现实任务、两种协作结构和三种协作模式的多模态具身协作基准测试平台。通过在多种MLLM上开展广泛实验,我们总结出三个核心发现:(i) 协作通常能提升具身任务完成率,但其收益取决于协作增益与协调复杂度之间的平衡;(ii) 通信是实现协作优势的关键,而最佳协作模式依赖于团队规模与模型能力;(iii) 协作还能增强在噪声先验和探索条件下的鲁棒性。总体而言,MECoBench为理解多模态具身协作的机制与边界提供了系统化测试平台。代码与数据集已开源至https://github.com/q-i-n-g/MECoBench。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an evaluation platform spanning diverse real-world tasks, two cooperation structures, and three collaboration modes. Through extensive experiments across various MLLMs, we summarize three key findings: (i) Collaboration generally improves embodied task completion, but its benefits depend on balancing collaborative gains against coordination complexity. (ii) Communication is essential to collaboration gains, while the best collaboration mode depends on team size and model capability. (iii) Moreover, collaboration improves robustness under noisy priors and exploration conditions. Generally, MECoBench provides a systematic testbed for understanding the mechanisms and limits of multimodal embodied collaboration. Code and dataset are available at https://github.com/q-i-n-g/MECoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。