无需人工标注,让模型自动生成并验证评分标准,提升奖励模型可靠性。
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

- 通过二元偏好训练模型生成与验证评分标准对,实现协作式纠错。
- 在RM-Bench上比基准模型高6.5分,长文本评估胜率提升6.0个百分点。
- 适合需要低成本、可扩展奖励建模的AI评估与对齐研究者使用。
评分标准增强的验证通过明确的评估准则引导奖励模型,比单一模型验证更可靠。但现有方法依赖昂贵的评分标准标注,限制了可扩展性。我们发现评分标准生成易受合作失败影响:低质量标准会主动误导奖励模型而非帮助。受合作通信原则启发,我们提出协同但批判的奖励建模(C2)框架,仅基于二元偏好训练奖励模型与评分标准生成器协同工作。C2通过衡量每个评分标准使奖励模型向正确偏好偏移或偏离的程度,合成有益与有害的评分标准对。利用这些对比对,训练出能提出有益标准的协作式生成器和仅在认为标准有效时才作判断的批判性验证器。C2在相同二元偏好数据上优于推理型奖励模型,在RM-Bench上最高提升6.5分,AlpacaEval 2.0长度控制胜率提升6.0个百分点。无需外部评分标准标注,8B规模的奖励模型性能可媲美4倍大模型提供的评分标准。本工作表明,通过激发评分标准增强验证中的有意合作,可在可扩展的前提下显著提升奖励模型可信度。
原文摘要 · Abstract (English)
Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. However, most existing methods require costly rubric annotations, limiting scalability. Moreover, we find that rubric generation is vulnerable to a failure of cooperation; low-quality rubrics actively mislead reward models rather than help. Inspired by the principle of cooperative communication, we propose Cooperative yet Critical reward modeling (C2), a framework that significantly improves reward model judgments by having the reward model critically collaborate with a rubric generator trained solely from binary preferences. In C2, we synthesize helpful and misleading rubric pairs by measuring how each rubric shifts the reward model toward or away from the correct preference. Using these contrastive pairs, we train a cooperative rubric generator to propose helpful rubrics, and a critical verifier to assess rubric validity before making its judgment, following only rubrics it deems helpful at inference time. C2 outperforms reasoning reward models trained on the same binary preferences, with gains of up to 6.5 points on RM-Bench and 6.0 points length-controlled win rate on AlpacaEval 2.0. Without external rubric annotations, C2 enables an 8B reward model to match performance achieved with rubrics from a 4$\times$ larger model. Overall, our work demonstrates that eliciting deliberate cooperation in rubric-augmented verification makes reward models more trustworthy in a scalable way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。