arXiv:2605.08761cs.MAcs.LG2026-05被引 1

评测大模型在企业多角色协作中的表现,聚焦权限与流程约束。

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

论文配图:Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
图 1 · 摘自论文原文
  • 构建11个角色分工的多智能体系统,模拟真实企业工作流。
  • 现有模型在任务委派、状态传递等环节成功率不足50%。
  • 适合研究企业级智能体协作与安全控制的学者和开发者。

大型语言模型代理正被期待应用于企业环境,但现有企业基准大多评估拥有广泛工具访问权限的单代理,而多代理基准很少体现角色专业化、权限控制、有状态业务系统及基于策略的审批等现实约束。本文提出 extsc{EntCollabBench},一个用于评估企业多代理协作的基准。该基准模拟一个由11个角色专业化代理组成的权限隔离组织,覆盖六个部门,包含两个评估子集:工作流子集要求代理协同修改企业系统状态,审批子集则要求代理做出基于政策的决策。评估基于执行轨迹、数据库状态验证和确定性策略裁决,而非自然语言响应判断。对代表性大模型代理的实验表明,当前模型在端到端企业协作中仍存在困难,尤其在任务委派、上下文传递、参数锚定、工作流闭环和决策承诺方面表现不佳。 extsc{EntCollabBench} 提供可复现的测试平台,用于衡量和改进面向真实组织环境的代理系统。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.

多智能体企业应用协作评测权限控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。