用合成数据和强化学习,让开源模型学会像闭源智能体一样高效搜信息。
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- 通过制造高不确定性任务,模拟复杂搜索场景来训练模型。
- 在BrowseComp测试中表现媲美闭源智能体,超越所有开源模型。
- 适合想提升搜索推理能力的研究者或开发者使用。
突破人类认知局限是大模型训练的关键前沿。像DeepResearch这样的闭源智能体在极复杂的资讯检索基准BrowseComp上展现出超人能力,此前开源模型无法企及。我们认为其成功源于一种开源模型缺乏的精细推理模式:在海量信息中系统性降低极端不确定性。基于此,我们提出WebSailor,一套完整的后训练方法,包括通过结构化采样与信息混淆生成新型高不确定性任务、RFT冷启动策略,以及高效的代理强化学习算法DUPO(Duplicating Sampling Policy Optimization)。通过这一集成流程,WebSailor在复杂信息检索任务中显著优于所有开源智能体,性能逼近闭源系统,有效弥合了能力差距。
原文摘要 · Abstract (English)
Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking benchmarks such as BrowseComp, a feat previously unattainable. We posit that their success hinges on a sophisticated reasoning pattern absent in open-source models: the ability to systematically reduce extreme uncertainty when navigating vast information landscapes. Based on this insight, we introduce WebSailor, a complete post-training methodology designed to instill this crucial capability. Our approach involves generating novel, high-uncertainty tasks through structured sampling and information obfuscation, RFT cold start, and an efficient agentic RL training algorithm, Duplicating Sampling Policy Optimization (DUPO). With this integrated pipeline, WebSailor significantly outperforms all open-source agents in complex information-seeking tasks, matching proprietary agents' performance and closing the capability gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。