arXiv:2606.03968cs.CLcs.AI2026-06

通过协同设计问题与评分标准,提升非可验证任务的强化学习效果。

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards

  • 用教师提炼的关键点重写开放式问题,生成可评估的场景化提问。
  • 基于教师策略差距构建对比评分标准,提升评估有效性。
  • 仅保留有信息量的问题-评分对训练,显著提升推理类任务表现。

基于评分标准的强化学习为拓展强化学习在非可验证奖励任务中的应用提供了新路径,但现有方法将查询分布视为固定,优化评分标准时忽视了其结构限制。开放性问题导致评分模糊;过度约束则引入模型无法验证的虚构参考,使所有响应均失效,训练失去奖励信号。本文提出QUBRIC框架,实现问题与评分标准的协同设计:利用教师提炼的关键点,将开放性问题重写为基于场景、可评估的问题;通过对比式评分生成,将教师策略差异转化为查询级评估标准;学习能力过滤机制保留具信息量的问答对用于GRPO训练。QUBRIC在ArenaHard上相较SFT基线提升5.5分。仅使用指令跟随数据训练后,该模型在法律、道德与叙事推理三个未见基准上平均提升6.3分,改进集中于推理维度。结果表明,协同设计问题与评分标准可使基于评分的强化学习成为超越严格可验证任务的实际补充。

原文摘要 · Abstract (English)

Rubric-based RL is a promising route for extending reinforcement learning beyond verifiable rewards, yet existing methods optimize rubrics while treating the query distribution as fixed. We identify a structural bottleneck: rubric quality is constrained by query structure. Open-ended queries yield vague rubrics; naively narrowing them introduces fabricated references that no model can verify, so all responses fail and training receives no reward signal. We present QUBRIC, a framework that co-designs queries and rubrics. Teacher-derived key points ground the rewriting of open-ended queries into scenario-based, evaluable questions. Contrastive rubric generation then turns teacher-policy gaps into query-level criteria, and learnability filtering retains only informative query-rubric pairs for GRPO training. QUBRIC achieves a +5.5 point gain on ArenaHard over the SFT baseline. Trained only on instruction-following data, it further transfers to three held-out benchmarks spanning legal, moral, and narrative reasoning (+6.3 points on average), with improvements concentrated in reasoning-related dimensions. These results provide evidence that co-designing queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks.

强化学习评分标准指令微调推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。