arXiv:2410.07553cs.AI2024-10被引 10

构建多智能体协作评估基准,测试语言沟通能力。

COMMA: A Communicative Multimodal Multi-Agent Benchmark

  • 设计多模态谜题任务,通过语言实现智能体间协作。
  • 顶级模型如GPT-4o在协作中表现不佳,部分模型低于随机基线。
  • 适合评估大模型在真实场景下跨智能体沟通的能力。

基于大模型的多模态智能体快速发展,但其在协作任务中基于语言的智能体间通信能力仍被忽视。现有评测基准未能充分覆盖智能体间通信与协作的关键方面,尤其在信息不对称、需联合完成个体无法独立完成的任务场景下。为此,我们提出COMMA:一个新型多模态多智能体协作评测基准,通过语言通信评估系统在四种核心能力上的表现。实验发现,包括GPT-4o和o4-mini在内的前沿模型存在显著缺陷,部分链式思维推理模型(如R1-Onevision、LLaVA-CoT)甚至无法超越随机基线,暴露出当前模型在智能体间沟通能力上的巨大提升空间。

原文摘要 · Abstract (English)

The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative tasks. This oversight presents a critical gap in understanding their effectiveness in real-world deployments, particularly when communicating with humans. Existing agentic benchmarks fail to address key aspects of inter-agent communication and collaboration, particularly in scenarios where agents have unequal access to information and must work together to achieve tasks beyond the scope of individual capabilities. To fill this gap, we introduce COMMA: a novel puzzle benchmark designed to evaluate the collaborative performance of multimodal multi-agent systems through language communication. Our benchmark features a variety of multimodal puzzles, providing a comprehensive evaluation across four key categories of agentic capability in a communicative collaboration setting. Our findings reveal surprising weaknesses in state-of-the-art models, including strong proprietary models like GPT-4o and reasoning models like o4-mini. Many chain of thought reasoning models such as R1-Onevision and LLaVA-CoT struggle to outperform even a random baseline in agent-agent collaboration, indicating a potential growth area in their communication abilities.

多智能体语言通信评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。