揭示智能体系统规模扩展的规律,指导不同任务选对协作方式。
Towards a Science of Scaling Agent Systems
- 通过260组实验建立可预测性能的量化模型。
- 多智能体协作在高单智能体表现时收益递减,工具任务有额外开销。
- 架构与任务不匹配会严重降低效果,适合优化智能体设计者。
基于语言模型的智能体系统在实际任务中广泛应用,但其性能随关键维度扩展的变化规律尚未被充分研究。本文提出一种定量扩展法则,作为预测模型,捕捉性能随协调机制、模型能力及可观测系统与任务因素的变化。我们在六个代理基准测试、五种典型架构(单智能体和四种多智能体:独立、集中式、去中心化、混合)以及三种LLM家族下,进行了260种配置的受控评估,标准化工具、提示和计算以隔离架构影响。所建模型在所有六项基准上达到交叉验证R²=0.373(使用任务驱动能力指标时为R²=0.413)。我们发现显著的能力饱和效应,并识别出三个模式:(1) 当单智能体基线性能达到一定水平后,协调带来的收益递减;(2) 工具密集型任务存在多智能体开销;(3) 无集中验证的架构更易传播错误。相对于单智能体基线,性能变化范围从分解型金融推理的+80.8%到顺序规划任务的-70.0%,表明架构与任务对齐决定协作成败。该框架在87%的未见配置中准确识别最优架构,并在前沿模型上保持一致的相对偏好。智能体效能取决于协调机制与任务结构的匹配度,不匹配将导致性能下降。
原文摘要 · Abstract (English)
Agents, language model-based systems capable of reasoning, planning, and acting are widely adopted in real-world tasks, yet how their performance changes as these systems scale across key dimensions remains underexplored. We introduce quantitative scaling principles for agent systems as a predictive model, capturing how performance varies with coordination, model capability, and measurable system and task factors. Across 260 configurations spanning six agentic benchmarks, five canonical architectures (Single-Agent and four Multi-Agent: Independent, Centralized, Decentralized, Hybrid), and three LLM families, we perform controlled evaluations, standardizing tools, prompts, and compute to isolate architectural effects. The resulting model achieves a cross-validated R^2=0.373 across all six benchmarks (R^2=0.413 with a task-grounded capability metric). We identify a robust capability-saturation effect and additional patterns: (1) a coordination yields diminishing returns once single-agent baselines exceed certain performance; (2) tool-heavy tasks appear to incur multi-agent overhead; and (3) architectures without centralized verification tend to propagate errors more than those with centralized coordination. Relative performance change compared to single-agent baseline ranges from +80.8% on decomposable financial reasoning to -70.0% on sequential planning, demonstrating that architecture-task alignment determines collaborative success. The framework identifies the best-performing architecture for 87% of held-out configurations and shows consistent relative architecture preferences on unseen frontier models. Agent effectiveness depends on alignment between coordination and task structure, and that mismatched coordination degrades the performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。