用多角色评估生成更全面的评分标准,提升大模型生成质量评判效果。
Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

- 通过多个互补角色生成评分维度,避免单一评估视角遗漏关键偏好。
- 在偏好验证任务中,比单角色方法在多个模型上均表现更优。
- 适合需要高可信度奖励信号的开放生成任务优化场景。
可靠的奖励与偏好信号对评估和优化大语言模型在开放式任务中的表现至关重要。基于评分表的评判方式能将判断分解为明确的评估标准,但现有无标注的评分表生成方法通常依赖单一通用评估者,易忽略人类偏好的重要维度,这种问题称为‘维度盲点’。为此,我们提出多角色评分表生成(MRRG)框架,无需训练、不依赖参考文本,通过调动多个互补角色获取评估标准,并整合成可审计的评分表评分器。该评分器可用于验证成对偏好,也可为基于可验证奖励的强化学习(RLVR)提供奖励信号。在偏好验证基准测试中,MRRG在多个主干模型上均持续优于单角色基线。进一步的RLVR实验表明,MRRG能生成更强的奖励信号,有效提升开放式生成能力。
原文摘要 · Abstract (English)
Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria, but existing annotation-free rubric generators typically rely on a single generic evaluator. As a result, they may overlook important dimensions of human preference, a failure mode we term dimensional blind spots. To address this limitation, we propose Multi-Role Rubric Generation (MRRG), a training-free and reference-free framework that elicits evaluation criteria from multiple complementary roles and consolidates them into an auditable rubric-based scorer. This scorer can be used both to validate pairwise preferences and to provide rewards for GRPO-style Reinforcement Learning with Verifiable Rewards (RLVR). Experiments on preference validation benchmarks show that MRRG consistently outperforms single-role rubric generation baselines across multiple backbone models. Further RLVR experiments demonstrate that MRRG yields a stronger reward signal for improving open-ended generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。