VRM让奖励模型更懂人类真实偏好,避免虚假关联。
VRM: Teaching Reward Models to Understand Authentic Human Preferences
- 用变分推断同时学习评价权重和语义特征,模拟人类判断过程。
- 在多个基准数据集上显著优于传统方法,更贴近真实人类评分。
- 适合关注对齐模型可靠性的研究者,尤其重视评估机制设计。
大型语言模型在多种自然语言任务中取得显著进展,但用于对齐的奖励模型常面临奖励黑客问题:其主要通过将提示-响应对直接映射为标量分数,可能捕捉到虚假相关性而非真实人类偏好。相比之下,人类评估涉及复杂过程:先根据提示上下文权衡多个高维目标的重要性,再通过低维语义特征(如逻辑连贯性、情境恰当性)评估响应质量。受此启发,我们提出VRM(变分奖励建模),一种新框架,通过将高维目标权重与低维语义特征作为隐变量,利用变分推断技术显式建模人类偏好判断过程。此外,我们提供了理论分析,表明VRM可实现比传统奖励模型更紧的泛化误差界。大量实验表明,VRM在多个基准数据集上显著优于现有方法,更能准确捕捉真实人类偏好。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable success across diverse natural language tasks, yet the reward models employed for aligning LLMs often encounter challenges of reward hacking, where the approaches predominantly rely on directly mapping prompt-response pairs to scalar scores, which may inadvertently capture spurious correlations rather than authentic human preferences. In contrast, human evaluation employs a sophisticated process that initially weighs the relative importance of multiple high-dimensional objectives according to the prompt context, subsequently evaluating response quality through low-dimensional semantic features such as logical coherence and contextual appropriateness. Motivated by this consideration, we propose VRM, i.e., Variational Reward Modeling, a novel framework that explicitly models the evaluation process of human preference judgments by incorporating both high-dimensional objective weights and low-dimensional semantic features as latent variables, which are inferred through variational inference techniques. Additionally, we provide a theoretical analysis showing that VRM can achieve a tighter generalization error bound compared to the traditional reward model. Extensive experiments on benchmark datasets demonstrate that VRM significantly outperforms existing methods in capturing authentic human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。