arXiv:2506.19262cs.CLcs.LG2025-06被引 2

适度多样性的生成数据能提升小样本模型性能,过度多样性反而有害。

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

  • 通过控制生成数据的多样性水平,研究其对下游模型的影响
  • 中等多样性生成数据在标签稀缺时显著提升模型表现
  • 适合关注数据质量与自生成数据训练策略的研究者

大型语言模型(LLMs)强大的生成能力使其成为缓解特定领域数据稀缺和减少人工标注成本的潜力方案。然而,近期研究表明,基于自生成数据进行迭代训练会导致模型崩溃,性能下降。尽管已有大量研究探讨自生成数据的影响,但往往忽视了数据多样性这一关键质量因素。本文旨在探究生成数据多样性对下游模型性能的影响。我们系统考察了不同多样性水平的生成数据对模型表现的影响,并分析了混合比例不同的生成数据(即合成数据)训练模型的效果。实验结果表明,在极小分布偏移条件下,适度多样性的生成数据可有效提升标签数据不足场景下的模型性能;而高度多样化的生成数据则产生负面影响。我们的实证发现为未来利用LLM作为数据生成器的研究提供了重要指导。

原文摘要 · Abstract (English)

With the remarkable generative capabilities of large language models (LLMs), using LLM-generated data to train downstream models has emerged as a promising approach to mitigate data scarcity in specific domains and reduce time-consuming annotations. However, recent studies have highlighted a critical issue: iterative training on self-generated data results in model collapse, where model performance degrades over time. Despite extensive research on the implications of LLM-generated data, these works often neglect the importance of data diversity, a key factor in data quality. In this work, we aim to understand the implications of the diversity of LLM-generated data on downstream model performance. Specifically, we explore how varying levels of diversity in LLM-generated data affect downstream model performance. Additionally, we investigate the performance of models trained on data that mixes different proportions of LLM-generated data, which we refer to as synthetic data. Our experimental results show that, with minimal distribution shift, moderately diverse LLM-generated data can enhance model performance in scenarios with insufficient labeled data, whereas highly diverse generated data has a negative impact. We hope our empirical findings will offer valuable guidance for future studies on LLMs as data generators.

生成数据多样性微调大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。