arXiv:2602.10531stat.MLcs.LG2026-02

模型训练中掺入合成数据反而能提升性能,关键在于保持真实数据的微量信息。

From Collapse to Improvement: Statistical Perspectives on the Evolutionary Dynamics of Iterative Training on Contaminated Sources

  • 从统计角度分析混合数据训练过程,揭示模型演化的动态机制。
  • 只要真实数据占比不为零,即使逐渐减少,也能避免性能崩溃并恢复目标分布。
  • 适用于语言模型等多类模型,对实际迭代训练有重要指导意义。

模型在合成数据上进行迭代训练时面临性能下降问题,本文从统计视角出发,指出只要存在来自真实目标分布的少量新鲜信息,即便数据被合成样本污染,仍可能实现性能提升。研究聚焦于真实分布与合成分布混合采样的迭代训练过程,在下一个词预测的语言模型中分析了整个演化路径,揭示混合权重与样本量如何共同控制长期性能。当真实分布具有非平凡的混合权重时,即使其随时间衰减,仅需以适当的样本量进行无感知污染的训练,即可避免模型崩溃,并在特定条件下恢复真实目标分布。模拟实验验证了结论,且该现象在其他模型类别中也具普遍性。

原文摘要 · Abstract (English)

The problem of model collapse has presented new challenges in iterative training of generative models, where such training with synthetic data leads to an overall degradation of performance. This paper looks at the problem from a statistical viewpoint, illustrating that one can actually hope for improvement when models are trained on data contaminated with synthetic samples, as long as there is some amount of fresh information from the true target distribution. In particular, we consider iterative training on samples sourced from a mixture of the true target and synthetic distributions. We analyze the entire iterative evolution in a next-token prediction language model, capturing how the interplay between the mixture weights and the sample size controls the overall long-term performance. With non-trivial mixture weight of the true distribution, even if it decays over time, simply training the model in a contamination-agnostic manner with appropriate sample sizes can avoid collapse and even recover the true target distribution under certain conditions. Simulation studies support our findings and also show that such behavior is more general for other classes of models.

生成模型模型崩溃迭代训练统计分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。