arXiv:2510.24701cs.CLcs.AI2025-10被引 21

通义深研模型可自主完成复杂长周期研究任务,性能领先。

Tongyi DeepResearch Technical Report

  • 采用端到端训练框架,融合智能体中期与后期训练,提升推理与信息搜索能力。
  • 参数量305亿,每令牌激活33亿,多项研究基准测试达顶尖水平。
  • 全自动化数据生成流水线,无需人工标注,适合科研与深度信息挖掘场景。

我们提出通义深研(Tongyi DeepResearch),一个专为长周期、深度信息检索研究任务设计的智能体大语言模型。通过端到端训练框架结合智能体中期训练与后期训练,实现复杂任务中可扩展的推理与信息搜索能力。设计了完全自动化的高可扩展数据合成流水线,无需依赖昂贵的人工标注,支持所有训练阶段。通过为各阶段构建定制化环境,确保系统交互稳定一致。该模型总参数量达305亿,每令牌仅激活33亿参数,在多项智能体深度研究基准测试中表现优异,涵盖Humanity's Last Exam、BrowseComp、BrowseComp-ZH、WebWalkerQA、xbench-DeepSearch、FRAMES及xbench-DeepSearch-2510。模型、框架与完整解决方案已开源,以赋能社区。

原文摘要 · Abstract (English)

We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.

智能体深度研究大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。