arXiv:2410.04203cs.AI2024-10ICLR被引 22

整合主流偏好优化组件,统一框架提升模型对齐效果

RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization

  • 将7类偏好优化组件统一到一个目标函数中
  • 在多个数据集上超越现有DPO方法表现
  • 为算法设计与实际应用提供可参考的指导

近期众多偏好优化算法作为直接偏好优化(DPO)家族的扩展被提出。尽管这些方法成功实现了模型与人类偏好的对齐,但对其附加组件贡献的理解仍不充分。此外,缺乏公平且一致的对比,难以判断哪些组件真正提升了下游性能。本文提出RainbowPO,一个统一框架,通过将现有DPO方法的关键组件归纳为七大方向,并整合进单一连贯的目标函数中,强化了各组件的表现。大量实验表明,RainbowPO优于现有DPO变体。同时,本工作为新DPO方法的设计提供洞见,也助力实践者高效实现。

原文摘要 · Abstract (English)

Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations.

偏好优化统一框架DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。