构建9000题深度研究数据集,助力智能体解决复杂现实问题
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
- 基于开源多跳问答数据自动生成高质量研究任务
- 涵盖9000个三阶难度问题,含可验证答案与推理轨迹
- 提供开源训练框架,支持强化学习与大模型评估
深度研究智能体具备多步网络探索、精准检索和复杂问答能力。然而,其发展面临两大瓶颈:缺乏大规模、真实场景下具有挑战性的数据集,以及缺少可访问的开源数据合成与训练框架。为此,我们构建了专为深度研究场景设计的大型挑战性数据集DeepResearch-9K,通过低成本自主管道从开源多跳问答数据集生成。该数据集包含9000个问题,分为L1至L3三个难度等级,配有由Tongyi-DeepResearch-30B-A3B这一先进深度研究智能体生成的高质量搜索轨迹与推理链,并提供可验证答案。此外,我们开发了开源训练框架DeepResearch-R1,支持多轮网页交互、多种强化学习方法及不同奖励模型(如基于规则的结果奖励与大模型判别反馈)。实验证明,在DeepResearch-R1上训练的智能体在多个挑战性深度研究基准测试中达到当前最优表现。数据集已发布于https://huggingface.co/datasets/artillerywu/DeepResearch-9K,代码开源于https://github.com/Applied-Machine-Learning-Lab/DeepResearch-R1。
原文摘要 · Abstract (English)
Deep-research agents are capable of executing multi-step web exploration, targeted retrieval, and sophisticated question answering. Despite their powerful capabilities, deep-research agents face two critical bottlenecks: (1) the lack of large-scale, challenging datasets with real-world difficulty, and (2) the absence of accessible, open-source frameworks for data synthesis and agent training. To bridge these gaps, we first construct DeepResearch-9K, a large-scale challenging dataset specifically designed for deep-research scenarios built from open-source multi-hop question-answering (QA) datasets via a low-cost autonomous pipeline. Notably, it consists of (1) 9000 questions spanning three difficulty levels from L1 to L3 (2) high-quality search trajectories with reasoning chains from Tongyi-DeepResearch-30B-A3B, a state-of-the-art deep-research agent, and (3) verifiable answers. Furthermore, we develop an open-source training framework DeepResearch-R1 that supports (1) multi-turn web interactions, (2) different reinforcement learning (RL) approaches, and (3) different reward models such as rule-based outcome reward and LLM-as-judge feedback. Finally, empirical results demonstrate that agents trained on DeepResearch-9K under our DeepResearch-R1 achieve state-of-the-art results on challenging deep-research benchmarks. We release the DeepResearch-9K dataset on https://huggingface.co/datasets/artillerywu/DeepResearch-9K and the code of DeepResearch-R1 on https://github.com/Applied-Machine-Learning-Lab/DeepResearch-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。