arXiv:2410.19720cs.CLcs.AI2024-10NAACL被引 9

让大模型理解人类偏好的两个维度,提升对齐效果。

2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision

  • 用段落和维度双信号监督模型生成
  • 在多个基准上优于传统单维度方法
  • 适合需要精细控制输出质量的研究者

最近的直接偏好优化(DPO)进展显著提升了大语言模型与人类偏好的对齐,因其简单高效。然而,现有方法通常仅优化标量分数或排序奖励,忽略了人类偏好的多维特性。本文提出将偏好扩展至二维:段落与方面。我们构建了名为HelpSteer-2D的二维监督数据集:对响应按句子分段并评分,同时设计多个覆盖响应质量的评判维度。基于二维信号,我们提出2D-DPO框架,将整体目标分解为多段落、多方面的子目标。在多个主流基准上的实验表明,2D-DPO优于仅优化标量或一维偏好的方法。

原文摘要 · Abstract (English)

Recent advancements in Direct Preference Optimization (DPO) have significantly enhanced the alignment of Large Language Models (LLMs) with human preferences, owing to its simplicity and effectiveness. However, existing methods typically optimize a scalar score or ranking reward, thereby overlooking the multi-dimensional nature of human preferences. In this work, we propose to extend the preference of DPO to two dimensions: segments and aspects. We first introduce a 2D supervision dataset called HelpSteer-2D. For the segment dimension, we divide the response into sentences and assign scores to each segment. For the aspect dimension, we meticulously design several criteria covering the response quality rubrics. With the 2-dimensional signals as feedback, we develop a 2D-DPO framework, decomposing the overall objective into multi-segment and multi-aspect objectives. Extensive experiments on popular benchmarks demonstrate that 2D-DPO performs better than methods that optimize for scalar or 1-dimensional preferences.

偏好优化多维反馈大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。