让AI Agent像人类一样长期深度研究,突破记忆瓶颈
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- 将研究过程建模为可迭代的决策过程,动态更新报告并保持专注空间
- 在6个基准测试中超越现有顶尖系统,包括闭源商业模型
- 适合需要复杂推理与长程规划的研究型AI任务
近期深度研究系统已展示出AI代理从外部源自主发现并整合知识的潜力。本文提出WebResearcher,一个构建此类代理的新框架,包含两个核心组件:(1) WebResearcher,一种迭代式深度研究范式,将深度研究重构为马尔可夫决策过程,使代理定期将成果整合进动态演进的报告中,同时维持专注工作区,克服现有单上下文方法面临的上下文窒息和噪声污染问题;(2) WebFrontier,一个可扩展的数据合成引擎,通过工具增强的复杂度递增生成高质量训练数据,系统性构建介于被动知识召回与主动知识建构之间的研究任务。值得注意的是,我们发现该范式生成的训练数据显著提升了传统单上下文方法的工具使用能力。此外,该范式可通过并行思考自然扩展,支持多代理并发探索以获得更全面结论。在6个挑战性基准上的大量实验表明,WebResearcher达到最先进性能,甚至超越前沿闭源系统。
原文摘要 · Abstract (English)
Recent advances in deep-research systems have demonstrated the potential for AI agents to autonomously discover and synthesize knowledge from external sources. In this paper, we introduce WebResearcher, a novel framework for building such agents through two key components: (1) WebResearcher, an iterative deep-research paradigm that reformulates deep research as a Markov Decision Process, where agents periodically consolidate findings into evolving reports while maintaining focused workspaces, overcoming the context suffocation and noise contamination that plague existing mono-contextual approaches; and (2) WebFrontier, a scalable data synthesis engine that generates high-quality training data through tool-augmented complexity escalation, enabling systematic creation of research tasks that bridge the gap between passive knowledge recall and active knowledge construction. Notably, we find that the training data from our paradigm significantly enhances tool-use capabilities even for traditional mono-contextual methods. Furthermore, our paradigm naturally scales through parallel thinking, enabling concurrent multi-agent exploration for more comprehensive conclusions. Extensive experiments across 6 challenging benchmarks demonstrate that WebResearcher achieves state-of-the-art performance, even surpassing frontier proprietary systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。