arXiv:2507.17746cs.LGcs.AI2025-07被引 289

用评分标准当奖励信号,让大模型在复杂任务中表现更优。

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

  • 将评分标准转化为结构化奖励,用于强化学习训练。
  • 在医疗和科学任务上提升31%和7%的性能表现。
  • 适合需要多维度评价的场景,尤其小模型更受益。

基于可验证奖励的强化学习(RLVR)在数学和编程等有明确正确性信号的任务中表现良好。然而,将其扩展到真实世界推理任务面临挑战,因为评估依赖于复杂的多维度判断,而非二值正确性。近期,实例级评分标准被用于评估基准以捕捉此类判断,但其作为在线策略微调奖励信号的潜力尚未充分探索。本文提出「评分标准作为奖励」(RaR),一种基于评分标准反馈的在线强化学习方法,将RLVR扩展至非可验证领域。我们在医学和科学领域评估了多种将评分标准反馈聚合为奖励的策略。最佳的RaR变体在HealthBench上相对主流的基于李克特量表奖励的LLM-as-judge基线提升了高达31%,在GPQA-Diamond上提升了7%。结果表明,经RaR训练的策略能适应多样化的评估格式,在基于评分标准和选择题任务中均表现优异。此外,使用评分标准作为结构化奖励信号可提升小规模裁判模型的对齐效果,并降低不同裁判尺度间的性能波动。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than binary correctness. Instance-specific rubrics have recently been used in evaluation benchmarks to capture such judgments, but their potential as reward signals for on-policy post-training remains underexplored. We introduce $\textbf{Rubrics as Rewards}$ (RaR), an on-policy reinforcement learning method that extends RLVR beyond verifiable domains by using rubric-based feedback. Across both medical and science domains, we evaluate multiple strategies for aggregating rubric feedback into rewards. The best RaR variant achieves relative improvements of up to $31\%$ on HealthBench and $7\%$ on GPQA-Diamond over popular LLM-as-judge baselines that rely on direct Likert-based rewards. These results demonstrate that RaR-trained policies adapt well to diverse evaluation formats, performing strongly on both rubric-based and multiple-choice tasks. Moreover, we find that using rubrics as structured reward signals yields better alignment for smaller judges and reduces performance variance across judge scales.

强化学习评分标准大模型训练多维度评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。