arXiv:2606.15367cs.AIcs.CL2026-06被引 3

构建可执行长程科研任务的智能代理,突破传统搜索式训练局限。

S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents

论文配图:S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents
图 1 · 摘自论文原文
  • 融合封闭问答与开放探索,构建图结构驱动的科研任务框架
  • 在20个基准上超越同规模开源模型,逼近顶级闭源模型性能
  • 适合需要深度推理、报告生成与多文件理解的研究型任务

深度研究代理旨在通过长期规划、证据收集、推理和报告生成来解决复杂知识密集型任务。尽管近期搜索代理在信息检索和答案验证方面取得进展,但现有训练数据集仍以封闭式问答和信息定位为主,主要训练信息获取行为,而对证据整合、知识合成、规划、文件理解及结构化报告生成等核心科研能力覆盖不足。本文提出一种统一的轨迹构建范式,结合封闭式问答与开放式探索,包含图基任务建模、代理轨迹生成和多维轨迹验证,实现高质、可扩展的长链复杂推理、深度研究指令遵循、报告撰写、文件理解与技能使用等能力的合成。相比现有搜索导向数据集,本方法更强调知识合成、复杂推理与规划。S1-DeepResearch-32B在涵盖五大能力维度(复杂推理、指令遵循、报告生成、文件理解、技能使用)的20个基准上达到同规模开源模型最优表现,在多个挑战性深度研究基准上接近领先闭源模型水平。结果表明,联合建模信息获取、知识合成与规划行为对构建高效深研代理至关重要。

原文摘要 · Abstract (English)

Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and report generation. While recent progress in search agents has demonstrated strong capabilities in information retrieval and answer verification, most existing training datasets remain search-centric, focusing primarily on closed-ended question answering and information localization. As a result, they mainly train information-seeking behavior while providing limited coverage of key deep research capabilities, including evidence integration, knowledge synthesis, planning, file understanding, and structured report generation. In this work, we propose a unified trajectory construction paradigm for deep research agents that combines closed-ended QA and open-ended exploration. The proposed framework consists of graph-grounded task formulation, agentic trajectory rollout, and multi-dimensional trajectory verification, enabling scalable synthesis of high-quality agentic trajectories spanning long-chain complex reasoning, deep research instruction following, report writing, file understanding and generation, and skills usage. Compared with existing search-oriented datasets, our synthesized trajectories place greater emphasis on knowledge synthesis, complex reasoning, and planning. S1-DeepResearch-32B achieves state-of-the-art performance among open-source models of comparable scale across 20 benchmarks spanning five capability dimensions, including complex reasoning, instruction following, report generation, file understanding, and skills usage. On several challenging deep research benchmarks, it approaches the performance of leading proprietary frontier models. These results highlight the importance of jointly modeling information acquisition, knowledge synthesis, and planning-oriented agent behaviors for building effective deep research agents.

科研代理长程推理知识合成多模态任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。