用多智能体生成研究级数据,让开源大模型变强
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
- 多智能体协作模拟工具使用,自动生成高质量研究数据
- 在多个规模模型上达成新SOTA,超越现有开源模型表现
- 无需私有数据,适合想提升开源模型的研究者
开源大语言模型与闭源模型之间的性能差距,主要源于高质量训练数据的获取差异。为缩小这一差距,我们提出一种自动化合成复杂、研究级指令数据的新框架。该方法基于多智能体工作流,通过协作型AI代理模拟集成工具的复杂推理过程,实现从头到尾的高保真数据生成。利用这些合成数据,我们设计了一种两阶段训练策略,结合监督微调与一种新型强化学习方法,以最大化模型对齐性与能力。大量实验证明,该框架可赋能多个规模的开源模型,在主流深度研究基准上达到新的最先进水平。本工作为不依赖专有数据或模型的开源大模型发展提供了一条可扩展、高效的道路。
原文摘要 · Abstract (English)
The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this gap, we introduce a novel framework for the automated synthesis of sophisticated, research-grade instructional data. Our approach centers on a multi-agent workflow where collaborative AI agents simulate complex tool-integrated reasoning to generate diverse and high-fidelity data end-to-end. Leveraging this synthesized data, we develop a two-stage training strategy that integrates supervised fine-tuning with a novel reinforcement learning method, designed to maximize model alignment and capability. Extensive experiments demonstrate that our framework empowers open-source models across multiple scales, enabling them to achieve new state-of-the-art performance on the major deep research benchmark. This work provides a scalable and effective pathway for advancing open-source LLMs without relying on proprietary data or models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。