用低成本训练让中等模型具备专家级深度研究能力
Step-DeepResearch Technical Report
- 基于原子能力的数据合成策略提升规划与报告生成
- 32B模型在Scale AI评测中达61.4%得分,超越多数开源模型
- 专为中文场景设计的ADR-Bench评测基准,填补评估空白
随着大语言模型向自主代理演进,深度研究能力成为关键指标。现有学术基准如BrowseComp难以满足开放性研究的真实需求,缺乏意图识别、长程决策和跨源验证等核心能力。为此,我们提出Step-DeepResearch,一种成本可控的端到端智能体。通过基于原子能力的数据合成策略强化规划与报告写作,并采用从代理中训练到SFT再到RL的渐进式训练路径。结合检查清单式评估器,显著提升系统鲁棒性。为弥补中文领域评估空白,我们构建了ADR-Bench,用于真实深度研究场景的评测。实验表明,Step-DeepResearch(32B)在Scale AI Research Rubrics上取得61.4%的分数;在ADR-Bench上显著优于同类开源模型,媲美OpenAI与Gemini DeepResearch等闭源领先模型。结果证明,精细化训练可使中等规模模型以行业领先的成本效率实现专家级表现。
原文摘要 · Abstract (English)
As LLMs shift toward autonomous agents, Deep Research has emerged as a pivotal metric. However, existing academic benchmarks like BrowseComp often fail to meet real-world demands for open-ended research, which requires robust skills in intent recognition, long-horizon decision-making, and cross-source verification. To address this, we introduce Step-DeepResearch, a cost-effective, end-to-end agent. We propose a Data Synthesis Strategy Based on Atomic Capabilities to reinforce planning and report writing, combined with a progressive training path from agentic mid-training to SFT and RL. Enhanced by a Checklist-style Judger, this approach significantly improves robustness. Furthermore, to bridge the evaluation gap in the Chinese domain, we establish ADR-Bench for realistic deep research scenarios. Experimental results show that Step-DeepResearch (32B) scores 61.4% on Scale AI Research Rubrics. On ADR-Bench, it significantly outperforms comparable models and rivals SOTA closed-source models like OpenAI and Gemini DeepResearch. These findings prove that refined training enables medium-sized models to achieve expert-level capabilities at industry-leading cost-efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。