arXiv:2510.07743cs.CL2025-10ACL被引 102

构建可扩展的合成评分标准库,提升大模型对齐效果

OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment

  • 通过对比优劣回答生成包含显式规则和隐含原则的评分标准
  • 在多个基准上比基线奖励模型性能提升8.4%
  • 适合需要高质量人工反馈对齐的LLM研究者使用

奖励建模是基于人类反馈的强化学习(RLHF)的核心,但现有奖励模型多依赖标量或成对判断,难以捕捉人类偏好的多维特性。近期研究提出将评分标准作为奖励(RaR),用结构化维度衡量回复质量。然而,生成可靠且可扩展的评分标准仍是关键挑战。本文提出OpenRubrics,一个大规模、多样化的(提示,评分标准)配对数据集,用于训练评分标准生成与基于评分标准的奖励模型。为获取区分性强且全面的评估信号,引入对比式评分标准生成(CRG),通过对比优选与拒选回答,提炼出硬性规则(显式约束)和原则(隐含品质)。进一步通过保持偏好标签一致性来剔除噪声评分标准。在多个奖励建模基准上,基于评分标准的奖励模型Rubric-RM相比同规模强基线提升8.4%。该提升可迁移至指令遵循和生物医学任务中的策略模型。

原文摘要 · Abstract (English)

Reward modeling lies at the core of reinforcement learning from human feedback (RLHF), yet most existing reward models rely on scalar or pairwise judgments that fail to capture the multifaceted nature of human preferences. Recent studies have explored rubrics-as-rewards (RaR) that uses structured criteria to capture multiple dimensions of response quality. However, producing rubrics that are both reliable and scalable remains a key challenge. In this work, we introduce OpenRubrics, a diverse, large-scale collection of (prompt, rubric) pairs for training rubric-generation and rubric-based reward models. To elicit discriminative and comprehensive evaluation signals, we introduce Contrastive Rubric Generation (CRG), which derives both hard rules (explicit constraints) and principles (implicit qualities) by contrasting preferred and rejected responses. We further remove noisy rubrics via preserving preference-label consistency. Across multiple reward-modeling benchmarks, our rubric-based reward model, Rubric-RM, surpasses strong size-matched baselines by 8.4%. These gains transfer to policy models on instruction-following and biomedical benchmarks.

奖励建模评分标准大模型对齐RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。