提出可解释的通用奖励模型,让大模型评估更透明可靠。
R3: Robust Rubric-Agnostic Reward Models
- 不依赖具体评分标准,跨任务通用
- 输出带推理过程,分数可解释
- 适合需要灵活对齐人类价值的场景
奖励模型在对齐语言模型输出与人类偏好方面至关重要,但现有方法往往缺乏可控性和可解释性。这些模型通常针对特定目标优化,限制了其在更广泛下游任务中的泛化能力。此外,其标量输出难以在无上下文推理的情况下理解。为此,我们提出了R3,一种新型奖励建模框架,具备评分标准无关性、跨评估维度的泛化能力,并提供可解释的推理式得分分配。R3支持更透明和灵活的语言模型评估,实现对多样化人类价值观和应用场景的稳健对齐。相关模型、数据与代码已开源,详见https://github.com/rubricreward/r3。
原文摘要 · Abstract (English)
Reward models are essential for aligning language model outputs with human preferences, yet existing approaches often lack both controllability and interpretability. These models are typically optimized for narrow objectives, limiting their generalizability to broader downstream tasks. Moreover, their scalar outputs are difficult to interpret without contextual reasoning. To address these limitations, we introduce $\shortmethodname$, a novel reward modeling framework that is rubric-agnostic, generalizable across evaluation dimensions, and provides interpretable, reasoned score assignments. $\shortmethodname$ enables more transparent and flexible evaluation of language models, supporting robust alignment with diverse human values and use cases. Our models, data, and code are available as open source at https://github.com/rubricreward/r3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。