发现大模型协作能力远低于单打独斗,提出新评测基准揭示合作短板。
The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation
- 设计无固定协议的迷宫协作评测,支持自然语言自由沟通
- 32个模型测试显示:单机表现强的模型协作时性能大幅下降
- 通过强模型引导弱模型可显著缩小协作差距,适合多智能体系统研究者
AI发展正趋向于由语言模型驱动的、由独立开发的异构智能体组成的系统,这些智能体拥有不同的信息、权限和工具。此类系统的成功依赖于异构智能体间在部分可观测条件下的有效协作。尽管关注度高,现有文献缺乏不依赖固定通信协议的实证研究,限制了对开放世界部署的理解。本文提出一个可示范的协作迷宫求解基准,具备:(i) 独立评估协作能力,(ii) 可调节问题复杂度,(iii) 支持可扩展自动化评分,(iv) 无输出格式约束,实现无引导、自然的通信评测。使用该基准,我们评估了32个领先开源与闭源模型在单人、同质及异质配对下的表现。结果揭示了一个令人意外的协作差距:在单人任务中表现优异的模型,在协作时性能显著下降。我们识别出若干缓解策略,例如采用‘接力推理’方法——由更强模型先行决策后移交较弱模型,可大幅缩小差距。研究主张:(1) 开展协作感知的评估,(2) 设计提升协作能力的训练策略,(3) 精心设计交互以激发智能体的协作潜力,这些原则适用于AI-AI及人机协同场景。
原文摘要 · Abstract (English)
The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools. The success of these systems will depend critically on effective collaboration among these heterogeneous agents, even under partial observability. Despite intense interest, the literature lacks empirical studies evaluating agentic collaboration without relying on fixed communication protocols, limiting insights for open-world deployments. We propose an illustrative collaborative maze-solving benchmark that (i) isolates collaborative capabilities, (ii) modulates problem complexity, (iii) enables scalable automated grading, and (iv) imposes no output-format constraints, benchmarking unguided, natural communication. Using this benchmark, we evaluate 32 leading open- and closed-source models in solo, homogeneous, and heterogeneous pairings. Our results reveal a surprising collaboration gap: models that perform well solo often degrade substantially when required to collaborate. We identify mitigations that show remarkable influence. For example, a small nudge via a relay inference approach, where a stronger agent leads before handing off to a weaker one, closes much of the gap. Our findings argue for (1) collaboration-aware evaluation, (2) training strategies to enhance collaborative capabilities, and (3) deliberate interaction design to elicit agents' collaboration skills, principles relevant to AI-AI and human-AI settings where agents must establish common ground.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。