用对比分析生成可解释的评分标准,提升大模型评价可靠性。
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
- 通过多维度对比学习识别评价差异的关键因素
- 仅用3000样本训练,性能超过全量微调模型
- 适合需要透明、高效奖励建模的研究与应用
奖励建模对对齐大语言模型与人类偏好至关重要,但传统方法可解释性差且依赖昂贵专家标注。现有基于评分标准的方法缺乏质量控制,导致标准冗余噪声,并未能缓解大模型评价中普遍存在的冗长性、位置等偏差,存在可扩展性与可靠性之间的权衡。为此,我们提出CDRRM(对比驱动的评分标准生成框架),采用‘对比-合成’新范式,先对偏好样本对进行多维对比分析,识别因果性判别因子,再将其整合为简洁、上下文感知的评分标准,指导偏好判断。在RewardBench、RMBench和RMB三个权威基准上的实验证明,CDRRM在多领域均达到领先性能,有效缓解了前述评价偏差。尤为关键的是,仅用3000个高质量样本训练评分生成器,即可使冻结的预训练判别模型超越全量微调基线。该工作为奖励建模提供了可扩展、可解释且数据高效的路径。
原文摘要 · Abstract (English)
Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on costly expert annotations. While recent rubric-based approaches enhance evaluation transparency, they lack systematic quality control, yielding noisy and redundant criteria, failing to mitigate persistent biases (e.g., verbosity, position) in LLM evaluators, and creating a scalability-reliability trade-off. To address these limitations, we propose CDRRM (Contrast-Driven Rubric Reward Model), a framework built on a novel Contrast-then-Synthesis paradigm for high-quality rubric generation and guided preference judgment. CDRRM first conducts multi-dimensional contrastive profiling on preference pairs to identify causal discriminative factors, then synthesizes these insights into compact, context-aware rubrics to guide preference judg- ments. Extensive experiments on three authoritative benchmarks (RewardBench, RMBench, RMB) demonstrate that CDRRM achieves state-of-the-art performance across diverse domains and effectively mitigates aforementioned evaluation biases. Notably, our approach delivers exceptional data efficiency: training the rubric generator on only 3k high-quality samples empowers a frozen pre-trained judge model to outperform fully fine-tuned baselines. This work offers a scalable, interpretable, and data-efficient path for reward modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。