arXiv:2603.00349cs.AIcs.MA2026-03

提出可量化协作的评估框架,让大模型多智能体系统协作过程可观察可修复。

COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems

  • 用任务进展定义协作要求,将语言交互与环境动作对齐
  • 在双环境三结构下提升任务成功率和约束满足率
  • 适合研究多智能体协作机制或需可靠执行的系统设计者

许多复杂任务需要持续努力、多样化能力或协调行动,单个智能体难以胜任。但简单增加智能体数量并不保证性能提升,因有效协作依赖于智能体间及与任务结构的动态互动。现有评估多将协作视为最终成功隐含因素,难以追踪协作过程与任务演化关系。本文提出COOP²框架,将大模型多智能体系统(LLM-MAS)的高层协作动态与环境任务进展相联系。该框架定义可验证协作要求的任务,分析协作随任务进展的演变,识别协作失效位置与原因。基于此,我们开发COOP²-Repair,通过群体计划预测约束失败,并开启定向修复通道。在两个环境和三种通信结构中,该方法提升了任务成功率与约束满足度,同时揭示了修复带来的额外决策开销与通信负担。

原文摘要 · Abstract (English)

Many complex tasks require extended effort, diverse capabilities, or coordinated actions beyond what a single agent can provide. However, simply adding more agents does not guarantee better performance, as effective cooperation depends on how agents interact with each other and with task structure to satisfy evolving constraints over time. This challenge is amplified for LLM-based multi-agent systems (LLM-MAS): plans, messages, and revisions occur in natural language, whereas task progress depends on grounded environment actions. Current evaluations mostly treat cooperation as an implicit ingredient of final task success, leaving both cooperation and the effect of multi-agent interaction on task dynamics difficult to study. We introduce COOP$^2$, an evaluation framework that grounds high-level agent cooperation dynamics in LLM-MAS within task progress in the environment. COOP$^2$ then defines cooperative tasks with verifiable cooperative requirements, allowing us to analyze how cooperation unfolds over time with respect to task progress, as well as where and why cooperation breaks down. Building on this framework, we develop COOP$^2$-Repair, which predicts constraint failures from group plans and opens targeted repair channels for guided revisions. Across two environments and three communication structures, COOP$^2$-Repair improves task success and constraint satisfaction while exposing the additional decision overhead and communication load required for repair. The project web page can be found at: https://happyeureka.github.io/coop2.

多智能体协作评估大模型任务修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。