AI编程助手协作时效率反降30%,因缺乏社会智能。
CooperBench: Why Coding Agents Cannot be Your Teammates Yet
- 设计600+真实开源任务,测试双代理协作能力
- 协同成功率比单独执行低30%,远逊人类团队
- 揭示沟通混乱、失信承诺、误判他人等三大缺陷
解决团队冲突不仅需要任务能力,还需社会智能以达成共识。随着AI代理参与复杂协作,其协调能力成为有效队友的关键。我们提出CooperBench基准,涵盖12个库、4种语言的600多个协作编码任务。每项任务分配两个代理独立实现功能,但可能产生冲突。任务基于真实开源仓库,配有专家编写测试。评估前沿编码代理发现:协同成功率平均比单独执行低30%,与人类团队显著相反。分析揭示三大问题:(1) 沟通渠道充斥模糊、时机不当和错误信息;(2) 即便沟通有效,代理仍会违背承诺;(3) 代理常对他人计划和沟通存在错误预期。大规模模拟中观察到角色分工、资源分配和协商等罕见涌现行为。本研究提供新型协作编码基准,呼吁从追求个体能力转向发展社会智能。
原文摘要 · Abstract (English)
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. As AI agents increasingly collaborate on complex work, they must develop coordination capabilities to function as effective teammates. Yet we hypothesize that current agents lack these capabilities. To test this, we introduce CooperBench, a benchmark of over 600 collaborative coding tasks across 12 libraries in 4 programming languages. Each task assigns two agents different features that can be implemented independently but may conflict without proper coordination. Tasks are grounded in real open-source repositories with expert-written tests. Evaluating state-of-the-art coding agents, we observe the curse of coordination: agents achieve on average 30% lower success rates when working together compared to performing both tasks individually. This contrasts sharply with human teams, where adding teammates typically improves productivity. Our analysis reveals three key issues: (1) communication channels become jammed with vague, ill-timed, and inaccurate messages; (2) even with effective communication, agents deviate from their commitments; and (3) agents often hold incorrect expectations about others' plans and communication. Through large-scale simulation, we also observe rare but interesting emergent coordination behavior including role division, resource division, and negotiation. Our research presents a novel benchmark for collaborative coding and calls for a shift from pursuing individual agent capability to developing social intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。