融合在线与离线数据,平衡分布偏移与质量,提升偏好优化效果
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
- 通过动态融合在线与离线数据,实现分布一致性与数据质量的权衡
- 在Arena-Hard上达到60.8%胜率,超越纯在线或离线数据方法
- 适合追求高精度偏好对齐的模型训练者,尤其适用于资源受限场景
直接偏好优化(DPO)通过人类偏好对齐语言模型。由于在线数据由策略模型直接生成,其分布与模型一致,通常表现更优。本文指出候选偏好样本的质量是另一关键因素:在线数据质量受策略模型能力限制,而离线数据虽存在分布偏移,但来源多样,潜力更大。然而,现有研究多依赖在线数据,忽视了离线数据在质量上的优势。为此,本文提出InCo-DPO,一种高效合成偏好数据的方法,可动态融合在线与离线数据,在分布偏移与数据质量间实现最优平衡。实验表明,该方法克服了离线数据的分布偏移问题和在线数据的质量瓶颈。在Alpaca-Eval 2.0与Arena-Hard基准上评估,InCo-DPO优于纯在线或离线数据方法,并在使用Gemma-2模型的vanilla DPO下取得Arena-Hard 60.8%的领先胜率。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) optimizes language models to align with human preferences. Utilizing on-policy samples, generated directly by the policy model, typically results in better performance due to its distribution consistency with the model compared to off-policy samples. This paper identifies the quality of candidate preference samples as another critical factor. While the quality of on-policy data is inherently constrained by the capabilities of the policy model, off-policy data, which can be derived from diverse sources, offers greater potential for quality despite experiencing distribution shifts. However, current research mostly relies on on-policy data and neglects the value of off-policy data in terms of data quality, due to the challenge posed by distribution shift. In this paper, we propose InCo-DPO, an efficient method for synthesizing preference data by integrating on-policy and off-policy data, allowing dynamic adjustments to balance distribution shifts and data quality, thus finding an optimal trade-off. Consequently, InCo-DPO overcomes the limitations of distribution shifts in off-policy data and the quality constraints of on-policy data. We evaluated InCo-DPO with the Alpaca-Eval 2.0 and Arena-Hard benchmarks. Experimental results demonstrate that our approach not only outperforms both on-policy and off-policy data but also achieves a state-of-the-art win rate of 60.8 on Arena-Hard with the vanilla DPO using Gemma-2 model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。