用深度研究构建可自生成的评分标准,提升开放任务奖励信号质量。
Deep Research as Rubric for Reinforcement Learning

- 通过多轮智能体搜索挖掘领域知识与失败模式,生成动态评分规则
- 仅需1K–3K样本即达强性能,自举版本在第三轮迭代表现最佳
- 适合需要细粒度奖励信号的智能体研究与专家推理任务
开放式推理与长文本生成任务缺乏可靠的自动验证信号用于基于奖励的策略优化。评分标准提供了一种有前景的替代方案,但现有方法将评分标准视为既定的静态模板——无论是人工设计或提示生成——常忽略任务特定、知识密集的关键维度,扭曲了奖励信号。我们的核心观察是:评分标准的构建本身就是一个研究问题:判断回答正确或深刻,需发现并整合外部知识。我们提出深度研究作为评分标准(DR-rubric),一种两阶段框架来构建此类评分标准。第一阶段通过多轮迭代式智能体搜索,提取领域事实、结构约束和失败模式;第二阶段将这些证据提炼为原子级、可独立验证的约束条件,供基于GRPO的策略优化使用。由于训练中的模型可充当自身评分标准生成器,DR-rubric-8B支持无需前沿模型辅助的自举式评分标准生成。我们在6个涵盖智能体研究与专家推理的任务上进行评估。实验表明,DR-Rubric仅需1K–3K训练实例即取得强竞争力表现:由GPT-5生成的评分标准在智能体任务中尤其提升广度覆盖,Gemini生成的评分标准在智能体与专家推理任务间表现最均衡,而自举评分标准展现出从专业化到再平衡的演进过程,在第三轮迭代时达到最优整体性能。结果证明,将评分标准构建从静态评估模板重构为基于证据的研究过程,能为开放任务提供更可扩展、更精细的奖励信号。
原文摘要 · Abstract (English)
Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but existing approaches treat them as given artifacts -- either hand-crafted or prompt-generated -- and often miss the task-specific, knowledge-intensive dimensions that matter most, distorting the reward signal. Our key observation is that rubric construction is itself a research problem: identifying what makes a response correct or insightful requires discovering and synthesizing external knowledge. We propose Deep Research as Rubric (DR-rubric), a two-stage framework for constructing such rubrics. Stage I elicits domain facts, structural constraints, and failure modes through iterative multi-turn agentic search; Stage II distills this evidence into atomic, independently verifiable constraints for GRPO-based policy optimization. Because the model under training can serve as its own rubric generator, DR-rubric-8B supports bootstrap rubric generation without frontier-model assistance. We evaluate on 6 benchmarks spanning agentic research and expert reasoning. Experiments show that DR-Rubric achieves strong competitive performance with only 1K -- 3K training instances, where GPT-5-generated rubrics particularly benefit breadth coverage on agentic tasks, Gemini-generated rubrics yield the most balanced performance across agentic and expert reasoning tasks, and bootstrap rubrics exhibit a specialization-to-rebalancing evolution achieving the best overall performance at the third iteration. Results demonstrate that reframing rubric construction from static evaluation templates into an evidence-driven research process yields more scalable, fine-grained reward signals for open-ended tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。