arXiv:2603.20380cs.MAcs.AI2026-03中稿 · HAXD 2026, 8 pages…被引 1

为可组合多智能体团队设计统一工具管理框架,提升协作效率与可控性。

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams

  • 提出CAT数据层,用纯文本文件统一定义智能体工具访问与行为
  • 在115项任务中验证22个模型(0.6B-35B参数)表现,完成超2500次执行
  • 适合需要高效协同的智能体系统开发者与工程化研究者

工业界和学术界研究人员常使用多智能体系统加速工作,但现有应用缺乏可扩展的智能体编排组件统一管理机制,影响人机交互质量及上下文工程协调能力。智能体行为规范分散于自然语言说明文件或框架内部配置,难以共享、版本化或跨团队协作维护。借鉴辐射安全中的ALARA原则(合理可行尽量低),本文提出上下文-智能体-工具(CAT)数据层,通过相互关联的纯文本文件,使用户能直接声明每个智能体的工具权限,并修改其使用的工具。我们通过命令行工具npcsh加载团队并执行智能体运行,评估了22个本地部署模型(参数量0.6B至35B)在115项实际任务上的表现,涵盖文件操作、网络搜索、多步脚本、工具链式调用及多智能体委派,共完成约2500次执行。分析了不同模型家族在各类任务中的成功与失效情况。

原文摘要 · Abstract (English)

Industry practitioners and academic researchers regularly use multi-agent systems to accelerate their work, but the applications through which users operate these systems do not provide a simple, unified mechanism for scalably managing critical components of the agent harness. This lack of control adversely impacts both the quality of individual human-agent interactions and reduces the capacity for practitioners to coordinate context engineering efforts. The behavioral specifications that define what agents in such systems can do remain fragmented across prose instruction files -- for which compliance cannot be guaranteed -- or framework-internal configurations, making these specifications difficult to share, version, or collaboratively maintain across teams and projects. Applying the ALARA principle from radiation safety (exposures kept as low as reasonably achievable) to context, we introduce a context-agent-tool (CAT) data layer expressed through interrelated plain-text files, allowing users to directly declare tool access for each agent and to modify the tools themselves that are used by the agents when processing. We demonstrate capability of this CAT data layer to enable real agentic usage by using a command-line shell that loads the team and executes agent runs -- \texttt{npcsh} -- and evaluating 22 locally-hosted models from 0.6B to 35B parameters across 115 practical tasks spanning file operations, web search, multi-step scripting, tool chaining, and multi-agent delegation. We characterize which model families succeed in certain task categories and where they break down across $\sim$2500 total executions.

多智能体工具管理可组合系统智能体工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。