arXiv:2602.01011cs.MAcs.AI2026-02中稿 · ICML被引 13

自组织大模型团队难以发挥专家优势,性能最高下降41.1%

Multi-Agent Teams Hold Experts Back

  • 让大模型自由协作,不设固定角色或流程
  • 专家表现被平均化,导致性能最高损失41.1%
  • 团队越大越倾向妥协,适合研究协作机制的学者

多智能体大模型系统正作为自主协作者部署,其成员自由交互而非遵循预设流程。在此类场景中,协调需通过互动自发形成。然而,现有工作多依赖固定角色、流程或聚合规则,未考察无约束自组织团队的实际表现。基于组织心理学,我们探究自组织团队能否实现强协同——即整体表现不低于或超过最优个体。在人类启发与前沿机器学习基准上,我们发现:与人类团队不同,大模型团队始终无法达到专家代理的性能,即使明确告知谁是专家,机器学习基准上性能损失高达41.1%。分解失败原因显示,关键瓶颈在于专家利用而非识别。对话分析揭示团队趋向整合性妥协——平均专家与非专家观点,而非合理加权,且该行为随团队规模增大而加剧,并与性能负相关。有趣的是,这种共识倾向提升了对对抗性代理的鲁棒性,暗示对齐与有效利用专家能力之间存在权衡。研究揭示了自组织多智能体团队在调动集体专长方面的显著差距。

原文摘要 · Abstract (English)

Multi-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, effective coordination cannot be fully designed in advance and must instead emerge through interaction. However, most prior work enforces coordination through fixed roles, workflows, or aggregation rules, leaving open the question of how well self-organizing teams perform when coordination is unconstrained. Drawing on organizational psychology, we study whether self-organizing LLM teams achieve strong synergy, where team performance matches or exceeds the best individual member. Across human-inspired and frontier ML benchmarks, we find that -- unlike human teams -- LLM teams consistently fail to match their expert agent's performance, even when explicitly told who the expert is, incurring performance losses of up to 41.1% on ML benchmarks. Decomposing this failure, we show that expert leveraging, rather than identification, is the primary bottleneck. Conversational analysis reveals a tendency toward integrative compromise -- averaging expert and non-expert views rather than appropriately weighting expertise -- which increases with team size and correlates negatively with performance. Interestingly, this consensus-seeking behavior improves robustness to adversarial agents, suggesting a trade-off between alignment and effective expertise utilization. Our findings reveal a significant gap in the ability of self-organizing multi-agent teams to harness the collective expertise of their members.

多智能体协作专家利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。