arXiv:2602.14069cs.CL2026-02被引 18

用可解释的评分体系替代模糊奖励,提升开放任务对齐的可靠性

Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric

  • 基于成对自适应评分标准动态生成可解释的评判规则
  • 在开放任务中实现比传统方法更优的判别力和抗欺骗能力
  • 适合需要透明、可调试对齐机制的研究者与应用

标量奖励模型将多维度人类偏好压缩为单一不透明分数,造成信息瓶颈,导致开放任务中易出现脆弱性和奖励作弊。我们主张,非可验证任务的稳健对齐本质上是原则泛化问题:奖励不应是内化的学习函数,而应是可检视的显式推理过程。为此,我们提出开放评分系统(OpenRS),一个基于评分标准的LLM裁判框架,核心为成对自适应元评分标准(PAMR)和轻量级逐项可验证评分标准(PVR),在有真实答案或程序化验证时提供硬约束和可验证奖励。OpenRS采用显式元评分标准——类似宪法的规范,控制评分标准的生成、加权与执行;通过条件化两个候选回答的语义差异,实时生成自适应评分标准;进行逐准则成对比较,并外部聚合准则级偏好,避免点对点加权标量融合,提升开放场景下的判别力。为保持跨领域的原则一致性与可编辑性,引入双层元评分标准优化流程(自动化演化优化通用原则,人类参与优化领域原则),并辅以可验证评分标准作为防止退化行为的护栏和客观子任务的可验证奖励来源。最终,我们将OpenRS应用于成对强化学习训练中的奖励监督。

原文摘要 · Abstract (English)

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking in open-ended alignment. We argue that robust alignment for non-verifiable tasks is fundamentally a principle generalization problem: reward should not be a learned function internalized into a judge, but an explicit reasoning process executed under inspectable principles. To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR) and lightweight Pointwise Verifiable Rubrics (PVRs), which provide both hard-constraint guardrails and verifiable reward components when ground-truth or programmatic checks are available. OpenRS uses an explicit meta-rubric -- a constitution-like specification that governs how rubrics are instantiated, weighted, and enforced -- and instantiates adaptive rubrics on the fly by conditioning on the semantic differences between two candidate responses. It then performs criterion-wise pairwise comparisons and aggregates criterion-level preferences externally, avoiding pointwise weighted scalarization while improving discriminability in open-ended settings. To keep principles consistent yet editable across various domains, we introduce a two-level meta-rubric refinement pipeline (automated evolutionary refinement for general principles and a reproducible human-in-the-loop procedure for domain principles), complemented with pointwise verifiable rubrics that act as both guardrails against degenerate behaviors and a source of verifiable reward for objective sub-tasks. Finally, we instantiate OpenRS as reward supervision in pairwise RL training.

强化学习可解释奖励对齐机制LLM裁判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。