用评分标准指导智能体分阶段完成复杂研究任务,提升学习效率。
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

- 以自生成评分标准分解研究任务,实现分阶段规划与执行
- 通过阶段化反馈机制,在4个长文本研究基准上超越开源模型
- 适合需要长期推理与可复用经验的复杂任务研究者
训练深度研究智能体——即能规划、搜索证据、评估并生成长篇报告的系统——已超出传统可验证奖励的强化学习范畴。其输出无真实答案,轨迹涉及多工具决策,且训练后缺乏将过往尝试转化为可复用经验的机制。本文主张评分标准不应仅作为最终答案评判工具,而应成为结构化策略执行、裁判反馈与智能体记忆的共享接口。基于此,提出RubricEM框架,结合分阶段策略分解与基于反思的元策略演化。RubricEM通过自生成评分标准使研究轨迹具备阶段感知能力,分别指导规划、证据收集、审查与合成。采用阶段化结构GRPO进行信用分配,利用各阶段评分判断提供更密集的语义反馈,支持长时序优化。同时,训练共享主干的反思型元策略,将被评估轨迹提炼为未来可用的评分引导。RubricEM-8B在四个长文本研究基准上表现优异,超越同类开源模型,接近专有深度研究系统。此外,通过深入分析揭示了RubricEM的关键构成要素。
原文摘要 · Abstract (English)
Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth answers, their trajectories span many tool-augmented decisions, and standard post-training offers little mechanism for turning past attempts into reusable experience. In this work, we argue that rubrics should serve not merely as final-answer evaluators, but as the shared interface that structures policy execution, judge feedback, and agent memory. Based on this view, we introduce RubricEM, a rubric-guided reinforcement learning framework that combines stagewise policy decomposition with reflection-based meta-policy evolution. RubricEM first makes research trajectories stage-aware by conditioning planning, evidence gathering, review, and synthesis on self-generated rubrics. It then assigns credit with Stage-Structured GRPO, which uses stagewise rubric judgments to provide denser semantic feedback for long-horizon optimization. In parallel, RubricEM trains a shared-backbone reflection meta-policy that distills judged trajectories into reusable rubric-grounded guidance for future attempts. The resulting RubricEM-8B achieves strong performance across four long-form research benchmarks, outperforming comparable open models and approaching proprietary deep-research systems. Beyond final performance, we perform thorough analyses to understand the key ingredients of RubricEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。