用可控实验框架评估大模型代理的协作能力
CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

- 基于人机协同理论构建可配置仿真框架
- 四款大模型测试显示协作表现受设计影响
- 适合研究多智能体系统协作机制的学者
基于大语言模型的多智能体系统展现出巨大潜力,其有效性依赖于智能体通过文本通道进行协作的能力,类似人类团队。然而,最新研究表明,多智能体系统失败并非因个体任务解决能力不足,而是缺乏协作能力:即建立共同基础、维持共享任务理解、平衡个体与集体利益、在互动中修复偏差的能力。计算机支持的协同工作(CSCW)领域数十年研究已明确人类团队在有限沟通下的协作要求,但现有评估仍主要关注任务结果或单个智能体的推理、规划与工具使用能力。为系统分析多智能体系统中智能体的协作能力,我们提出CollabSim——一个融合理论驱动的协作能力定义、交互条件的可控操控以及智能体内部状态的动作级探测的可配置仿真框架。在四个大语言模型上的实验表明,CollabSim能够捕捉条件效应,区分模型性能模式,并揭示智能体设计的任务依赖性影响。
原文摘要 · Abstract (English)
Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate through text-based channels much as human teams do. Yet recent study suggests that MAS often falter not because agents lack individual task-solving ability, but because they lack collaborative competence: the capacity to establish common ground, maintain shared task understanding, balance individual and collective incentives, and repair misalignment as interaction unfolds. Decades of research in Computer-Supported Cooperative Work have characterized these requirements for human teams coordinating under constrained communication, yet existing MAS evaluations focus mainly on task outcomes or single-agent proficiency in reasoning, planning, and tool use. To enable a systematic analysis of agents' collaborative competence in MAS, we introduce CollabSim, a configurable simulation framework that combines a theory-grounded definition of collaborative capabilities, controlled manipulation of interaction conditions, and action-level probing of agents' internal states. Experiments across four LLMs show that CollabSim can capture condition effects, separate model performance patterns, and reveal task-dependent effects of agent design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。