arXiv:2506.06522cs.CLcs.AI2025-06NeurIPS被引 7

对比两个开源大模型后训练数据集,优化出更高效的新数据配方。

Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance

  • 用细粒度标注分析数据结构与质量差异
  • 新数据混合版TuluTalk减少14%样本但性能更优
  • 适合想高效构建高质量训练数据的研究者

近期大语言模型(LLM)的后训练与对齐研究越来越依赖于精心设计的数据集,以提升指令遵循、世界知识和专业技能。然而,多数领先开源与闭源模型所用的后训练数据仍不公开,其构建过程缺乏透明度。这一现状促使了开源后训练语料库的发展。尽管在这些开放数据上训练可达到与领先模型相当的性能,但由于大规模严谨比较所需计算成本过高,系统性评估仍极为困难,因此相关研究几乎空白。这导致我们尚不清楚特定样本、任务类型或数据筛选策略如何影响下游表现。本文首次对两个主流开源后训练数据集——Tulu-3-SFT-Mix 和 SmolTalk——进行全面对比分析。基于 Magpie 框架,我们为每条样本标注了转录结构(单轮/多轮)、任务类别、输入质量与输出质量等详细指标,并统计出二者在结构与质量上的异同。据此,我们设计了一套有原则的数据筛选方案,生成新数据混合体 TuluTalk:其样本量比任一原始数据集减少14%,但在关键基准测试中表现持平或超越原数据集。研究结果为在资源约束下构建更有效的后训练数据提供了可操作的指导。为支持未来研究,我们公开发布原始标注数据集及经优化的 TuluTalk 数据混合体。

原文摘要 · Abstract (English)

Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible to the public, with limited information about their construction process. This lack of transparency has motivated the recent development of open-source post-training corpora. While training on these open alternatives can yield performance comparable to that of leading models, systematic comparisons remain challenging due to the significant computational cost of conducting them rigorously at scale, and are therefore largely absent. As a result, it remains unclear how specific samples, task types, or curation strategies influence downstream performance when assessing data quality. In this work, we conduct the first comprehensive side-by-side analysis of two prominent open post-training datasets: Tulu-3-SFT-Mix and SmolTalk. Using the Magpie framework, we annotate each sample with detailed quality metrics, including turn structure (single-turn vs. multi-turn), task category, input quality, and response quality, and we derive statistics that reveal structural and qualitative similarities and differences between the two datasets. Based on these insights, we design a principled curation recipe that produces a new data mixture, TuluTalk, which contains 14% fewer samples than either source dataset while matching or exceeding their performance on key benchmarks. Our findings offer actionable insights for constructing more effective post-training datasets that improve model performance within practical resource limits. To support future research, we publicly release both the annotated source datasets and our curated TuluTalk mixture.

大模型训练数据质量后训练数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。