arXiv:2511.10985cs.CLcs.AI2025-11被引 1

首个系统分析开源偏好数据集,构建更优更小的混合数据集UltraMix。

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets

  • 用奖励模型自动标注偏好数据,实现可扩展的质量评估。
  • 新数据集UltraMix比最优单个数据集小30%但性能更优。
  • 适合关注数据质量与模型对齐效率的研究者。

大语言模型对齐是后训练的核心目标,常通过奖励建模与强化学习实现。其中,直接偏好优化(DPO)因在偏好补全与非偏好补全间微调模型而广受采用。尽管主流大模型未公开其偏好对,社区已发布TuluDPO、ORPO、UltraFeedback、HelpSteer和Code-Preference-Pairs等开源DPO数据集。然而系统性比较稀缺,主要因计算成本高且缺乏丰富质量标注,难以理解偏好选择机制、任务类型覆盖及每样本的人类判断一致性。本文首次对主流开源DPO语料库进行数据驱动的全面分析。我们利用Magpie框架为每个样本标注任务类别、输入质量与偏好奖励(基于奖励模型的信号,无需人工标注),实现可扩展、细粒度的偏好质量检查,揭示了各数据集中奖励差距的结构性与定性差异。基于此,我们系统性地构建新混合数据集UltraMix,从五个数据集中选择性提取并剔除噪声或冗余样本。UltraMix比表现最佳的单一数据集小30%,却在关键基准上超越其性能。我们公开所有标注、元数据与所构建的混合数据集,以推动数据驱动的偏好优化研究。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct preference optimization (DPO) has emerged as a widely adopted technique that fine-tunes LLMs on preferred completions over less favorable ones. While most frontier LLMs do not disclose their curated preference pairs, the broader LLM community has released several open-source DPO datasets, including TuluDPO, ORPO, UltraFeedback, HelpSteer, and Code-Preference-Pairs. However, systematic comparisons remain scarce, largely due to the high computational cost and the lack of rich quality annotations, making it difficult to understand how preferences were selected, which task types they span, and how well they reflect human judgment on a per-sample level. In this work, we present the first comprehensive, data-centric analysis of popular open-source DPO corpora. We leverage the Magpie framework to annotate each sample for task category, input quality, and preference reward, a reward-model-based signal that validates the preference order without relying on human annotations. This enables a scalable, fine-grained inspection of preference quality across datasets, revealing structural and qualitative discrepancies in reward margins. Building on these insights, we systematically curate a new DPO mixture, UltraMix, that draws selectively from all five corpora while removing noisy or redundant samples. UltraMix is 30% smaller than the best-performing individual dataset yet exceeds its performance across key benchmarks. We publicly release all annotations, metadata, and our curated mixture to facilitate future research in data-centric preference optimization.

偏好优化数据集分析模型对齐DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。