评测大模型在多智能体协调提示中的表现,发现多数模型能力有限。
PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

- 构建110个场景的基准,测试模型生成多智能体协作提示的能力。
- 平均通过率仅17.2%,顶级模型GPT-5.5达62.0%,信息泄露率高达217.9%。
- 适合研究多智能体系统、提示工程与大模型协同能力的学者和工程师。
真实世界的大模型应用正从单智能体流程转向多智能体协同系统,但当前模型仍难以确定各子智能体所需知识。为此,我们提出PerspectiveGap基准,用于评估大模型在多智能体系统中生成协调提示的能力。该基准包含110个场景,采用两种掺杂干扰项的任务形式:角色片段分配与自由格式提示撰写。这些场景按10种拓扑结构组织,源于作者的实际工程经验,并基于提示经济原则:构建以循环为中心的协调机制,以最小的角色和工程开销实现最大效用。在33个来自10家公司的商用模型测试中,GPT-5.5显著优于其他模型,而Opus 4.8尽管编码能力强,但在协调提示任务中表现不佳。总体上,模型平均综合通过率仅为17.2%(GPT-5.5为62.0%),平均总泄漏率高达217.9%(每场景信息泄露事件数,非比例;GPT-5.5为49.1%)。结果表明,多智能体协调提示是一项独立且未被充分评估的能力,PerspectiveGap为此提供了系统化测量与改进的基础。
原文摘要 · Abstract (English)
Real-world LLM applications are moving beyond single-agent workflows toward orchestrated multi-agent systems, yet current models still struggle to determine what each sub-agent needs to know. To measure this, we introduce PerspectiveGap, a benchmark for evaluating LLMs' ability to compose orchestration prompts for multi-agent systems. PerspectiveGap contains 110 scenarios, each evaluated through two distractor-mixed task formats: role-fragment assignment and free-form prompt writing. These scenarios are organized into 10 topologies, which are distilled from the authors' real-world engineering practice and framed by the Prompt Economy principle: building loop-centered orchestrations that maximize utility with minimal role and engineering overhead. In experiments with 33 commercial models from 10 companies, GPT-5.5 substantially outperforms all competitors, whereas Opus 4.8 shows a notable weakness in orchestration prompting despite its strong coding performance. Nevertheless, PerspectiveGap remains challenging: the evaluated models achieve an average combined pass rate of only 17.2\% (GPT-5.5 62.0\%) and an average overall leakage rate of 217.9\% (a per-scenario information leak-event count, not a proportion; GPT-5.5 49.1\%). These findings suggest that multi-agent orchestration prompting is a distinct and under-evaluated capability, and PerspectiveGap provides a foundation for measuring and improving it systematically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。