提出蜜蜂式架构,让企业智能体在受控环境中安全执行复杂任务。
Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration
- 用女王蜂控制平面生成结构化指令,由特化工蜂代理执行。
- 任务成功率96.4%,零治理失败,执行范围更精准。
- 适合需要权限隔离与审计的大型企业场景。
企业智能体系统需连接大模型与私有工具、内部知识及模型上下文协议(MCP)接口。单纯任务能力不足:组织还需策略执行、租户级隔离及明确操作边界内的执行。我们提出Queen-Bee架构,其女王控制平面检索能力、规划任务范围执行,并生成由受限工具访问的专用蜜蜂代理执行的结构化BeeSpec。实现包含租户级MCP连接器、基于审计的运行时治理、检索驱动的弱孵化和多种资源配置后端的原型系统。在59个企业级任务上评估,涵盖治理敏感请求、检索驱动配置、范围化本地执行和化学工作流集成。检索驱动的Queen-Bee变体达成0.964的任务成功率、零治理失败,且执行质量显著优于静态基准与宽松单代理基准。进一步展示具备显式审批门控的多蜜蜂化学工作流,以及基于真实上游证据和筛选产物的前3名推荐清单。与混合检索和LLM引导配置的对比显示,更丰富的配置后端虽可行,但在当前小规模、高度结构化的能力注册表中未超越轻量级结构化检索器。结果提供原型级系统证据,而非生产部署研究,提示企业智能体平台应不仅评估能力,还需考察受控配置、隔离行为、范围执行质量与产物感知的工作流协调。
原文摘要 · Abstract (English)
Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) interfaces. In this setting, raw task capability is insufficient: organizations also require policy enforcement, tenant-scoped isolation, and execution that remains within explicit operational boundaries. We present Queen-Bee, a governed multi-agent architecture in which a Queen control plane retrieves capabilities, plans task-scoped execution, and compiles a structured BeeSpec that is executed by specialized Bee agents under constrained tool access. We implement a working prototype with tenant-scoped MCP connectors, audit-backed execution-time governance, retrieval-driven weak incubation, and multiple provisioning backends. We evaluate the system on 59 enterprise-style tasks spanning governance-sensitive requests, retrieval-driven provisioning, scoped local execution, and chemistry workflow integration. The retrieval-driven Queen-Bee variant achieves a task success rate of 0.964, zero governance failures, and substantially better scoped execution quality than both a static Queen-Bee baseline and a permissive single-agent baseline. We further show a multi-Bee chemistry workflow with explicit approval gating and a concrete top-3 shortlist grounded in real upstream evidence and screening artifacts. Additional comparisons with hybrid retrieval and LLM-guided provisioning show that richer provisioning backends are viable but do not outperform the lightweight structured retriever on the current small, highly structured capability registry. The results provide prototype-level systems evidence rather than a production deployment study, and suggest that enterprise agent platforms should be evaluated not only by capability, but also by governed provisioning, isolation behavior, scoped execution quality, and artifact-aware workflow coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。