通过并行调用工具提升深度研究代理效率,显著减少推理步骤。
W&D:Scaling Parallel Tool Calling for Efficient Deep Research Agents
- 在单步推理中内建并行工具调用,无需复杂多智能体协作。
- 在BrowseComp上达62.2%准确率,比原GPT-5-High提升7.3个百分点。
- 适合追求高效深度推理的科研自动化场景,尤其关注响应速度。
深度研究代理通过多步推理和基于网络的信息检索,已成为自动化复杂智力任务的强大工具。尽管近期工作通过增加序列思维和工具调用次数提升了代理的深度,但通过并行工具调用扩展宽度的潜力仍未被充分探索。本文提出广深研究代理(Wide and Deep research agent)框架,旨在研究在增加深度的同时,通过并行工具调用扩展宽度对代理行为与性能的影响。不同于依赖复杂多智能体编排的方法,本方法利用内在的并行工具调用机制,在单个推理步骤内实现有效协调。实验表明,扩展宽度可显著提升深度研究基准表现,并减少获得正确答案所需的交互轮数。我们通过案例研究分析了提升因素,并探索了多种工具调用调度策略以优化并行策略。结果表明,权衡宽度与深度是构建高效深度研究代理的关键路径。值得注意的是,无需上下文管理等技巧,仅用GPT-5-Medium在BrowseComp上即达62.2%准确率,超越原GPT-5-High报告的54.9%。
原文摘要 · Abstract (English)
Deep research agents have emerged as powerful tools for automating complex intellectual tasks through multi-step reasoning and web-based information seeking. While recent efforts have successfully enhanced these agents by scaling depth through increasing the number of sequential thinking and tool calls, the potential of scaling width via parallel tool calling remains largely unexplored. In this work, we propose the Wide and Deep research agent, a framework designed to investigate the behavior and performance of agents when scaling not only depth but also width via parallel tool calling. Unlike existing approaches that rely on complex multi-agent orchestration to parallelize workloads, our method leverages intrinsic parallel tool calling to facilitate effective coordination within a single reasoning step. We demonstrate that scaling width significantly improves performance on deep research benchmarks while reducing the number of turns required to obtain correct answers. Furthermore, we analyze the factors driving these improvements through case studies and explore various tool call schedulers to optimize parallel tool calling strategy. Our findings suggest that optimizing the trade-off between width and depth is a critical pathway toward high-efficiency deep research agents. Notably, without context management or other tricks, we obtain 62.2% accuracy with GPT-5-Medium on BrowseComp, surpassing the original 54.9% reported by GPT-5-High.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。