新基准评估大模型奖励模型对用户偏好的个性化捕捉能力。
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- 基于用户特定规则构建优劣响应对,精准衡量个性化偏好。
- 现有顶级奖励模型个性化准确率仅75.94%,表现不佳。
- 该基准与下游任务表现相关性更强,适合真实场景评估。
多元对齐已成为大语言模型发展的重要前沿,奖励模型(RMs)是捕捉多样化人类价值观的核心机制。尽管通用响应质量评估基准广泛存在,但如何评估奖励模型对个体用户偏好的适应能力仍是一大挑战。为此,我们提出Personalized RewardBench,一个旨在严格评估奖励模型个性化建模能力的新基准。通过严格遵循(或违背)用户特定评分标准构建优选与次选响应对,确保偏好差异完全源于个人偏好,同时保持两者的通用质量(如正确性、相关性和帮助性)。人工评估证实,响应对之间的主要区分因素为个人偏好。大量测试显示,现有顶尖奖励模型在个性化任务上表现不佳,最高准确率仅为75.94%。关键的是,由于有效基准应能预测下游性能,我们验证发现该基准在Best-of-N采样和近端策略优化(PPO)中与下游表现的相关性显著高于现有基线。这些结果确立了Personalized RewardBench作为奖励模型下游应用性能的可靠代理工具。
原文摘要 · Abstract (English)
Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values. While benchmarks for general response quality are prevalent, evaluating how well reward models account for individual user preferences remains an open challenge. To bridge this gap, we introduce Personalized RewardBench, a novel benchmark designed to rigorously assess reward models' capacity to model personalized preferences. We construct chosen and rejected response pairs based on strict adherence to (or violation of) user-specific rubrics, ensuring that preference distinctions are uniquely tailored to the individual. In particular, human evaluations confirm that the primary discriminative factor between pairs is strictly personal preference, with both responses maintaining high general quality (e.g., correctness, relevance and helpfulness). Extensive testing reveals that existing state-of-the-art reward models struggle significantly with personalization, peaking at an accuracy of just 75.94%. Crucially, because an effective reward model benchmark should predict a reward model's performance on downstream tasks, we conduct experiments demonstrating that our benchmark exhibits a significantly higher correlation with downstream performance in both Best-of-N (BoN) sampling and Proximal Policy Optimization (PPO) compared to existing baselines. These findings establish Personalized RewardBench as a robust and accurate proxy for evaluating reward models' performance in downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。