arXiv:2605.24218cs.CL2026-05被引 6

用合成数据训练出能处理复杂研究任务的开源大模型。

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

论文配图:QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
图 1 · 摘自论文原文
  • 用统一评分树生成可验证奖励的合成数据,免人工标注。
  • 仅用8000个合成任务,性能超越多个闭源前沿模型。
  • 适合想构建通用研究型AI的研究者与开发者使用。

深度研究代理将搜索引擎从关键词匹配升级为知识合成,彻底改变人机信息交互方式。然而,前沿系统仍属私有,现有开源代理在不同任务类型间泛化能力差,缺乏有效训练方法。我们发布QUEST,一个涵盖2B至35B参数量的开源模型系列,作为通用深度研究代理,具备事实查询、引用定位和报告生成等强能力。训练采用中段监督微调与强化学习结合的方案,核心是基于统一评分树的高质量合成数据流水线,支持多任务类型且无需人工标注即可生成可验证奖励的数据。此外,QUEST内置上下文管理机制,支持长时程推理与知识整合。仅用8000个合成任务,QUEST在八个跨领域深度研究基准上达到或超过前沿闭源代理表现,成为近期开源模型中综合性能最佳者。全部模型、数据与训练脚本均已开源。

原文摘要 · Abstract (English)

Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how humans interact with information. However, frontier systems remain proprietary, while existing open agents often generalize poorly across different task types, leaving unclear how to train a broadly capable deep research agent. We release QUEST, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of long-horizon search tasks, with strong capabilities in fact seeking, citation grounding, and report synthesis. To build QUEST, we propose an effective training recipe combining mid-training, supervised fine-tuning, and reinforcement learning. Central to this recipe is a curated data synthesis pipeline based on unified rubric trees, which applies to different task types and enables synthesizing training data with verifiable rewards without human annotation. In addition, QUEST incorporates a built-in context management mechanism that enables effective long-horizon reasoning and knowledge synthesis. Using only 8K synthesized tasks, QUEST approaches or even surpasses frontier closed-source agents across eight deep research benchmarks spanning diverse task types, and achieves the best overall performance among recent open-weight agents. We released everything: models, data, and training scripts.

研究代理合成数据开源模型长程推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。