让开源模型具备超人类网页搜索推理能力
WebSailor: Navigating Super-human Reasoning for Web Agent

- 通过构造高不确定性任务,训练模型系统性降低复杂信息中的困惑度
- 在BrowseComp基准上达到与闭源模型相当的性能,超越所有开源代理
- 适合追求顶尖信息检索能力的研究者和开发者
突破人类认知局限是大语言模型训练的关键前沿。如DeepResearch等专有智能体已在极端复杂的资讯搜索基准BrowseComp上展现出超人类能力,此前开源模型无法企及。我们认为其成功源于开源模型所缺乏的一种精妙推理模式:在海量信息中系统性降低极端不确定性。基于此洞察,我们提出WebSailor——一种完整的后训练方法,旨在赋予模型这一核心能力。该方法包括通过结构化采样与信息混淆生成新型高不确定性任务、采用RFT冷启动策略,以及高效智能体强化学习算法DUPO(复制采样策略优化)。通过这一集成流程,WebSailor在复杂信息搜索任务中显著优于所有开源代理,性能逼近且匹配专有模型,有效弥合了能力差距。
原文摘要 · Abstract (English)
Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking benchmarks such as BrowseComp, a feat previously unattainable. We posit that their success hinges on a sophisticated reasoning pattern absent in open-source models: the ability to systematically reduce extreme uncertainty when navigating vast information landscapes. Based on this insight, we introduce WebSailor, a complete post-training methodology designed to instill this crucial capability. Our approach involves generating novel, high-uncertainty tasks through structured sampling and information obfuscation, RFT cold start, and an efficient agentic RL training algorithm, Duplicating Sampling Policy Optimization (DUPO). With this integrated pipeline, WebSailor significantly outperforms all opensource agents in complex information-seeking tasks, matching proprietary agents' performance and closing the capability gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。