提出多维度评估方法,揭示奖励模型在偏好理解上的短板。
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- 构建六项探测任务,从多个维度检验奖励模型能力
- 发现模型在多维度偏好上表现不佳,需优化多目标建模
- 引入推理时探测法,提升预测可信度与可解释性
以往奖励模型评估仅依赖固定成对排序测试集,缺乏对各偏好维度的表现分析。本文提出探测偏好表征的评估方法,构建多维度奖励模型基准(MRMBench),包含六项针对不同偏好维度的探测任务,旨在鼓励模型更好捕捉多维偏好。进一步提出推理时探测方法,可识别预测过程中使用的偏好维度,增强可解释性。大量实验表明,MRMBench与大语言模型对齐性能高度相关,是开发先进奖励模型的可靠参考。分析显示,当前奖励模型普遍难以同时把握多维偏好,凸显多目标优化的潜力。此外,推理时探测法能提供可靠的预测置信度指标,有助于提升语言模型对齐效果。
原文摘要 · Abstract (English)
Previous methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with the alignment performance of large language models (LLMs), making it a reliable reference for developing advanced reward models. Our analysis of MRMBench evaluation results reveals that reward models often struggle to capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Additionally, our findings show that the proposed inference-time probing method offers a reliable metric for assessing the confidence of reward predictions, which ultimately improves the alignment of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。