arXiv:2510.05133cs.CL2025-10

研究合成数据比例对模型性能的影响,发现30%是关键阈值。

Characterizing Model Behavior Under Synthetic Data Training: An Empirical Study Across Scales and Mixing Ratios

  • 控制合成数据比例,测试不同规模模型在多种任务上的表现
  • 合成数据超30%时性能急剧下降,大模型更抗干扰
  • 校准度先于准确率下降,可提前预警;适合模型训练调参者

由大语言模型生成的合成数据已成为现代自然语言处理训练流程的核心组成部分,从启动推理能力到扩充指令遵循数据集均有应用。尽管近期工作证明在保持高外部数据比例下可取得成功,但关于合成数据比例如何影响不同规模模型行为的系统性理解仍不充分。本文通过受控实证研究,考察了在不同合成数据与外部数据比例下模型的性能、校准性和输出特征。基于Pythia模型系列(410M–12B参数)在五个多样化任务上的实验,评估了模型在经过一至三次训练迭代、合成数据占比0%-50%的情况。主要发现包括:合成数据占比不超过20%时模型性能稳定,超过30%后退化加速;6.9B–12B的大模型比410M–1.4B的小模型更具鲁棒性;校准度下降早于准确率损失,可作为早期预警信号;任务特性显著影响退化速度,推理类任务比检索类任务退化更快。重要的是,当前最佳实践如STaR和Self-Instruct系统中保持外部数据比例高于80%,处于我们实验所识别的安全区间内。本文为从业者提供基于模型规模和任务需求的合成数据预算建议,并与同期研究(如Shumailov等人关于模型坍缩的发现)进行了详细对比。

原文摘要 · Abstract (English)

Synthetic data generated by large language models has become integral to modern NLP training pipelines, from bootstrapping reasoning capabilities to augmenting instruction-following datasets. While recent work demonstrates successful applications maintaining high external data ratios, systematic understanding of how synthetic data proportion affects model behavior across different scales remains limited. This paper presents a controlled empirical study examining model performance, calibration, and output characteristics when trained on varying synthetic-to-external data ratios. Using the Pythia model suite (410M-12B parameters) across five diverse tasks, we evaluate models after one to three training iterations with synthetic data proportions ranging from 0-50\%. Our key findings include: models maintain stable performance with up to 20\% synthetic data, but degradation accelerates beyond 30\%; larger models (6.9B-12B) show greater robustness to synthetic data than smaller models (410M-1.4B); calibration degradation precedes accuracy loss, providing an early warning signal; and task characteristics matter, with reasoning tasks degrading faster than retrieval tasks under synthetic data training. Importantly, we find that current best practices, such as those employed in STaR and Self-Instruct systems that maintain greater than 80\% external data, operate well within safe regimes identified by our experiments. We provide practical guidance for practitioners on synthetic data budgets based on model scale and task requirements, alongside detailed comparison with concurrent work including Shumailov et al.'s model collapse findings.

合成数据模型训练大模型实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。