arXiv:2505.02363cs.CL2025-05ICML被引 4

混合在线与离线数据,显著提升语言模型对齐效果

SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning

  • 简单混合在线与离线偏好数据,实现互补优势
  • 在Alpaca Eval 2.0上平均提升6.03%,优于单一数据源
  • 适合需要兼顾推理与创意任务的模型对齐场景

语言模型对齐人类偏好依赖成对偏好数据。尽管一些研究认为在线数据在偏好学习中始终优于离线数据,但也有研究表明其优势可能依赖具体任务,凸显系统探索二者交互关系的必要性。本文发现,在线数据在数学与编程等推理任务中表现更优,而离线数据在创作写作和个人推荐等开放任务中更具优势。基于此,我们提出SIMPLEMIX,通过简单混合两类数据,融合其互补优势。在多种任务与基准上的实证结果表明,SIMPLEMIX显著提升了语言模型对齐效果:相比在线DPO与离线DPO,平均在Alpaca Eval 2.0上提升6.03%;且优于更复杂的混合方法(如HyPO和DPO-Mix-P),平均提升3.05%。

原文摘要 · Abstract (English)

Aligning language models with human preferences relies on pairwise preference datasets. While some studies suggest that on-policy data consistently outperforms off -policy data for preference learning, others indicate that the advantages of on-policy data may be task-dependent, highlighting the need for a systematic exploration of their interplay. In this work, we show that on-policy and off-policy data offer complementary strengths in preference optimization: on-policy data is particularly effective for reasoning tasks like math and coding, while off-policy data performs better on open-ended tasks such as creative writing and making personal recommendations. Guided by these findings, we introduce SIMPLEMIX, an approach to combine the complementary strengths of on-policy and off-policy preference learning by simply mixing these two data sources. Our empirical results across diverse tasks and benchmarks demonstrate that SIMPLEMIX substantially improves language model alignment. Specifically, SIMPLEMIX improves upon on-policy DPO and off-policy DPO by an average of 6.03% on Alpaca Eval 2.0. Moreover, it outperforms prior approaches that are much more complex in combining on- and off-policy data, such as HyPO and DPO-Mix-P, by an average of 3.05%.

偏好学习数据混合语言模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。