仅用10.6k数据点,纯学术团队靠简单微调就训练出顶尖搜索代理。
OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories

- 用三类数据增强生成高难度探索轨迹,提升训练效率
- 仅10.6k样本即在4个基准上达领先水平,最高准确率78.0%
- 首次纯学术团队在同规模下超越工业级大模型,代码开源
深度搜索能力已成为前沿大语言模型智能体的关键能力,但其发展长期被工业巨头主导。典型工业流程需经历预训练、持续预训练(CPT)、监督微调(SFT)和强化学习(RL)等多阶段,资源消耗巨大。本文表明,若使用信息丰富且难度高的轨迹数据,仅通过简单的SFT方法即可实现强大性能。我们引入三项数据合成改进:扩大知识图谱规模以促进更丰富的探索,扩展工具集以增强功能覆盖,采用严格低步数过滤。基于仅10.6k数据点的训练,OpenSeeker-v2在四个基准测试中达到顶尖表现(30B规模、ReAct范式):浏览任务评测(BrowseComp)46.0%,中文版(BrowseComp-ZH)58.1%,人类终极考试(Humanity's Last Exam)34.6%,xbench 78.0%。该成绩超越了经过重资产CPT+SFT+RL训练的通义深研(Tongyi DeepResearch),后者分别为43.4%、46.7%、32.9%、75.0%。值得注意的是,OpenSeeker-v2是首个由纯学术团队使用仅SFT方法,在同等模型规模与范式下达到业界领先水平的搜索代理。我们已开源模型权重,共享这一简单而高效的方法,旨在使前沿搜索代理研究更具可及性。
原文摘要 · Abstract (English)
Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continual pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). In this report, we show that when fueled with informative and high-difficulty trajectories, a simple SFT approach could be surprisingly powerful for training frontier search agents. By introducing three simple data synthesis modifications: scaling knowledge graph size for richer exploration, expanding the tool set size for broader functionality, and strict low-step filtering, we establish a stronger baseline. Trained on merely 10.6k data points, our OpenSeeker-v2 achieves state-of-the-art performance across 4 benchmarks (30B-sized agents with ReAct paradigm): 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on Humanity's Last Exam, and 78.0% on xbench, surpassing even Tongyi DeepResearch trained with heavy CPT+SFT+RL pipeline, which achieves 43.4%, 46.7%, 32.9%, and 75.0%, respectively. Notably, OpenSeeker-v2 represents the first state-of-the-art search agent within its model scale and paradigm to be developed by a purely academic team using only SFT. We are excited to open-source the OpenSeeker-v2 model weights and share our simple yet effective findings to make frontier search agent research more accessible to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。