arXiv:2605.29156cs.LGcs.CL2026-05被引 2

用交替训练提升大模型在主观领域评分的准确性与稳定性。

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

  • 交替训练评分生成器与条件判别器,结合概率评分减少评分僵局。
  • 仅用成对偏好数据实现有效强化学习,保持评分一致性。
  • 适合需要高精度主观评估的非可验证领域任务如创意写作。

点对点奖励建模为大模型后训练提供关键信号,但在主观且不可验证的场景中面临绝对评分困难。基于评分表的方法通过分解评价标准缓解此问题,但现有方法通常依赖前沿大模型,并因硬性布尔聚合导致评分僵局。本文提出 RUBRIC-ARROW,一种交替框架,联合训练评分生成器与条件判别器,其强化学习阶段仅使用成对偏好数据。该方法结合基于概率的评分规则以降低僵局,并采用分阶段偏好奖励与交替 GRPO 策略,共同训练点对点评估器。大量实验表明,RUBRIC-ARROW 在奖励建模准确率上表现优异,并在下游策略后训练中持续带来增益。

原文摘要 · Abstract (English)

Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address this by decomposing evaluation into explicit criteria, but existing approaches typically depend on frontier LLMs and suffer from ties caused by hard Boolean aggregation. We present RUBRIC-ARROW, an alternating framework that jointly trains a rubric generator and a rubric-conditioned judge, with its RL stage using only pairwise preference data. Our method couples a probability-based scoring rule that reduces ties with phase-specific preference-based rewards and an alternating GRPO scheme that together train the pointwise evaluator. Extensive experiments show that RUBRIC-ARROW achieves competitive reward-modeling accuracy and yields consistent gains for downstream policy post-training.

奖励建模大模型强化学习评分表

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。