提出递归优化框架RRD,让大模型评分标准更全面、精准、无冗余。
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
- 用递归分解与过滤机制,拆分并净化评分标准
- 在JudgeBench上提升评分准确率最高达+17.7分
- 适合需要高质量评分信号的开放任务训练场景
近期,评分标准被用于引导大模型判断者捕捉主观、多维的人类偏好,并扩展至强化学习微调(RFT)中的奖励信号。然而,评分标准生成难以控制:常存在覆盖不足、维度混淆、偏好方向错位及冗余或高度相关指标,降低判断准确性并导致次优奖励。本文提出RRD框架,基于递归分解-过滤循环,将粗粒度标准分解为细粒度、可区分的判别准则,扩大覆盖范围并增强响应间区分度;辅以互补过滤机制剔除错位与冗余标准,并采用相关性感知加权方案避免高度相关准则过度代表,最终生成信息丰富、全面且非冗余的标准集。实证表明,RRD在评估与训练中均带来显著且一致的提升:在JudgeBench和PPE上,对GPT-4o与Llama3.1-405B判断者均提高偏好判断准确率,所有设置下表现最优,最高达+17.7分。作为WildChat的奖励源用于RFT时,产生更强更稳定的信号,使奖励提升最高达160%(Qwen3-4B)和60%(Llama3.1-8B),优于此前基线的10-20%;增益还能迁移至HealthBench-Hard与BiGGen Bench。总体而言,RRD确立了递归评分标准优化在开放领域大模型判断与奖励建模中的可扩展、可解释基础。
原文摘要 · Abstract (English)
Recently, rubrics have been used to guide LLM judges in capturing subjective, nuanced, multi-dimensional human preferences, and have been extended from evaluation to reward signals for reinforcement fine-tuning (RFT). However, rubric generation remains hard to control: rubrics often lack coverage, conflate dimensions, misalign preference direction, and contain redundant or highly correlated criteria, degrading judge accuracy and producing suboptimal rewards during RFT. We propose RRD, a principled framework for rubric refinement built on a recursive decompose-filter cycle. RRD decomposes coarse rubrics into fine-grained, discriminative criteria, expanding coverage while sharpening separation between responses. A complementary filtering mechanism removes misaligned and redundant rubrics, and a correlation-aware weighting scheme prevents over-representing highly correlated criteria, yielding rubric sets that are informative, comprehensive, and non-redundant. Empirically, RRD delivers large, consistent gains across both evaluation and training: it improves preference-judgment accuracy on JudgeBench and PPE for both GPT-4o and Llama3.1-405B judges, achieving top performance in all settings with up to +17.7 points on JudgeBench. When used as the reward source for RFT on WildChat, it yields substantially stronger and more stable learning signals, boosting reward by up to 160% (Qwen3-4B) and 60% (Llama3.1-8B) versus 10-20% for prior rubric baselines, with gains that transfer to HealthBench-Hard and BiGGen Bench. Overall, RRD establishes recursive rubric refinement as a scalable and interpretable foundation for LLM judging and reward modeling in open-ended domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。