arXiv:2410.15226cs.CL2024-10被引 67

提出新方法评估合成数据多样性,发现多样性显著影响大模型训练效果。

On the Diversity of Synthetic Data and its Impact on Training Large Language Models

  • 设计基于LLM聚类的多样性度量方法,量化合成数据差异性。
  • 实验表明多样性越高,预训练与微调性能越好,尤其微调阶段更敏感。
  • 适合关注数据生成效率与模型训练优化的研究者参考。

大型语言模型(LLMs)的发展凸显了对多样化、高质量预训练数据的需求。合成数据成为应对数据稀缺与获取困难的可行方案。以往研究多聚焦真实数据的质量与数量,而本文首次实现对合成数据多样性的可测量,并探究其对LLM性能的影响。通过引入新的多样性度量方法——LLM cluster-agent,我们系统考察了合成数据多样性在350M与1.4B参数模型的预训练及微调阶段的下游影响。一系列受控实验表明,该基于聚类的LLM评分方法所测得的多样性与预训练及监督微调性能呈正相关。研究还发现,预训练阶段的合成数据多样性对监督微调的影响,甚至超过对预训练本身的直接影响,这一现象在较小模型中同样显著。本研究深化了对合成数据在LLM训练中最优应用的理解,为高效数据生成提供了新方向。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has accentuated the need for diverse, high-quality pre-training data. Synthetic data emerges as a viable solution to the challenges of data scarcity and inaccessibility. While previous literature has focused predominantly on the quality and quantity of real data, our work enables the measurement of diversity in synthetic data and explores its impact on LLM performance. We study the downstream effects of synthetic data diversity during both the pre-training and fine-tuning stages by introducing a new diversity metric, \textit{LLM cluster-agent}, designed to evaluate the diversity of synthetic datasets. Through a series of controlled experiments with models of 350M and 1.4B parameters, we demonstrate that the proposed cluster-based LLM scoring of diversity correlates positively with both pre-training and supervised fine-tuning performance. Our findings also reveal that synthetic data diversity in pre-training affects supervised fine-tuning more significantly than pre-training itself, even for smaller models. We hope this study advances our understanding of the optimal use of synthetic data in LLM training and opens new avenues for efficient data generation processes.

合成数据大模型训练多样性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。