arXiv:2602.01511cs.CLcs.LG2026-02被引 42

用强化学习动态生成评分标准,让大模型更懂非可验证任务的质量评判。

Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training

  • 通过强化学习交替优化评分标准生成与判断模型
  • 在多个基准上超越现有方法,显著提升策略对齐效果
  • 适合需要高质量非可验证输出的场景,如创意写作

标准奖励模型通常预测单一分数,难以捕捉创造性写作或开放指令遵循等不可验证领域中响应质量的多维特性。为此,我们提出 Rubric-ARM 框架,通过偏好反馈的强化学习联合优化评分标准生成器与判别器。不同于依赖静态评分标准或分离训练流程的方法,本方法将评分标准生成视为可学习的隐式动作,以最大化判断准确性。我们引入交替优化策略,缓解同时更新带来的非平稳性问题,并提供理论分析证明该调度可降低训练过程中的梯度方差。大量实验表明,Rubric-ARM 在多个基准上达到当前最优性能,并在离线与在线强化学习设置下显著提升下游策略对齐效果。

原文摘要 · Abstract (English)

Standard reward models typically predict scalar scores that fail to capture the multifaceted nature of response quality in non-verifiable domains, such as creative writing or open-ended instruction following. To address this limitation, we propose Rubric-ARM, a framework that jointly optimizes a rubric generator and a judge using reinforcement learning from preference feedback. Unlike existing methods that rely on static rubrics or disjoint training pipelines, our approach treats rubric generation as a latent action learned to maximize judgment accuracy. We introduce an alternating optimization strategy to mitigate the non-stationarity of simultaneous updates, providing theoretical analysis that demonstrates how this schedule reduces gradient variance during training. Extensive experiments show that Rubric-ARM achieves state-of-the-art performance among baselines on multiple benchmarks and significantly improves downstream policy alignment in both offline and online reinforcement learning settings.

强化学习评分模型大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。