30B模型实现顶尖研究能力,靠三智能体协作与分阶段训练。
Mind DeepResearch Technical Report

- 三智能体协同:规划、搜索、报告分工,各司其职。
- 多项指标领先:在中文浏览、深度研究等任务上超同类开源模型。
- 真实场景验证:基于500个真实用户查询的评测,适合落地应用研究。
我们提出Mind DeepResearch(MindDR),一种高效的多智能体深度研究框架,仅用约300亿参数模型即达到领先性能,依托精心设计的数据合成与多阶段训练流程。其核心创新为协作式三智能体架构(规划、深度搜索、报告)及四阶段专业化训练流程:SFT冷启动、搜索强化学习、报告强化学习与偏好对齐。该方案使MindDR在~30B规模下表现优异,具体在BrowseComp-ZH达45.7%,BrowseComp达42.8%,WideSearch达46.5%,xbench-DS达75.0%,DeepResearch Bench达52.5%,超越同等规模开源系统,媲美更大模型。该系统已部署于理想汽车在线产品中。此外,我们构建了MindDR Bench,包含500个来自内部产品的真实中文查询,采用多维度评估体系而非单一RACE指标。在该基准上,MindDR取得51.8的最新最优成绩。
原文摘要 · Abstract (English)
We present Mind DeepResearch (MindDR), an efficient multi-agent deep research framework that achieves leading performance with only ~30B-parameter models through a meticulously designed data synthesis and multi-stage training pipeline. The core innovation of MindDR lies in a collaborative three-agent architecture (Planning Agent, DeepSearch Agent, and Report Agent) and a four-stage agent-specialized training pipeline comprising SFT cold-start, Search-RL, Report-RL and preference alignment. With this regime, MindDR demonstrates competitive performance even with ~30B-scale models. Specifically, MindDR achieves 45.7% on BrowseComp-ZH, 42.8% on BrowseComp, 46.5% on WideSearch, 75.0% on xbench-DS, and 52.5 on DeepResearch Bench, outperforming comparable-scale open-source agent systems and rivaling larger-scale models. MindDR has been deployed as an online product in Li Auto. Furthermore, we introduce MindDR Bench, a curated benchmark of 500 real-world Chinese queries from our internal product user interactions, evaluated through a comprehensive multi-dimensional rubric system rather than relying on a single RACE metric. On MindDR Bench, MindDR achieves a state-of-the-art score of 51.8.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。