arXiv:2511.19399cs.CLcs.AI2025-11被引 84

用动态评分标准训练出首个完全开源的长篇深度研究模型

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

  • 训练时让评分标准随模型演化,自动吸收新信息提升判断力
  • 在科学医疗等4个领域任务中平均比现有开源模型高15.6%
  • 性能媲美闭源模型但成本低1000倍,适合资源有限的研究者

深度研究代理需完成多步推理以生成长篇、有引证的答案。然而,多数开源深度研究代理通过强化学习在易验证的短问答任务上训练,难以扩展到真实长篇任务。本文提出基于动态评分的强化学习(RLER),让评分标准与策略模型共同演化,在训练中融入搜索结果和模型响应对比的新信息,实现更精准的事实核查和更具区分度的反馈。基于此方法,我们构建了首个直接针对开放式长篇深度研究训练的完全开源模型——Deep Research Tulu(DR Tulu-8B)。在科学、医疗及通用领域的四个长篇深度研究基准测试中,该模型平均比现有开源代理高出15.6%,与闭源代理相比仅低0.7%,同时单次查询成本仅为闭源模型的1/1000。

原文摘要 · Abstract (English)

Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via reinforcement learning with verifiable rewards, which does not extend to realistic long-form tasks. We address this with Reinforcement Learning with Evolving Rubrics (RLER), where rubrics are constructed and maintained to co-evolve with the policy model during training. This allows the rubrics to incorporate newly explored information from search and contrasting model responses, enabling better fact checking and more discriminative on-policy feedback. Using RLER, we develop Deep Research Tulu (DR Tulu-8B), the first fully open model that is directly trained for open-ended, long-form deep research. Across four long-form deep research benchmarks in science, healthcare, and general domains, DR Tulu substantially outperforms existing open deep research agents (by 15.6% over Tongyi DR on average) and matches or exceeds proprietary deep research agents (by 0.7% over OpenAI DR on average), while being significantly smaller and cheaper per query (1000x cheaper than OpenAI DR per query).

深度研究强化学习开源模型动态评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。