arXiv:2506.14157cs.CLcs.AI2025-06EMNLP被引 3

提出新指标DCRM,评估偏好数据对模型训练效果的影响

DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization

  • 用距离与奖励差值构建综合指标DCRM,衡量响应对质量
  • 高DCRM的训练集能显著提升模型在AlpacaEval等评测上的表现
  • 基于DCRM优选响应对,适用于改进大模型偏好优化数据集

近期研究尝试将偏好优化(PO)性能与底层偏好数据集关联。本文观察到,被偏好响应 $y^+$ 与不被偏好响应 $y^-$ 之间的差异影响大模型的学习能力,但这些差异未必符合理想学习需求。为此,我们引入距离与奖励边际来量化差异,并结合两者得到距离校准奖励边际(DCRM),用于衡量偏好优化中响应对的质量。直观上,DCRM鼓励最小化噪声差异、最大化期望差异。基于此,我们分析三类常见偏好数据集,按响应来源和偏好标注函数分为两类维度。实证发现,训练集的高DCRM与更好学习效果存在普遍相关性。受此启发,我们提出 best-of-$N^2$ 配对方法,选取DCRM最高的响应对。实验表明,在多种设置下,该方法生成的数据集可进一步提升模型在 AlpacaEval、MT-Bench 与 Arena-Hard 上的表现。

原文摘要 · Abstract (English)

Recent research has attempted to associate preference optimization (PO) performance with the underlying preference datasets. In this work, our observation is that the differences between the preferred response $y^+$ and dispreferred response $y^-$ influence what LLMs can learn, which may not match the desirable differences to learn. Therefore, we use distance and reward margin to quantify these differences, and combine them to get Distance Calibrated Reward Margin (DCRM), a metric that measures the quality of a response pair for PO. Intuitively, DCRM encourages minimal noisy differences and maximal desired differences. With this, we study 3 types of commonly used preference datasets, classified along two axes: the source of the responses and the preference labeling function. We establish a general correlation between higher DCRM of the training set and better learning outcome. Inspired by this, we propose a best-of-$N^2$ pairing method that selects response pairs with the highest DCRM. Empirically, in various settings, our method produces training datasets that can further improve models' performance on AlpacaEval, MT-Bench, and Arena-Hard over the existing training sets.

偏好优化数据质量评分指标LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。