arXiv:2505.13947stat.MLcs.LG2025-05被引 6

从概率视角揭示语言模型训练中崩溃现象的成因与缓解方法

A Probabilistic Perspective on Model Collapse

  • 将递归训练视为随机游走,分析样本量与估计偏差对模型演化的影响
  • 证明逐步增大样本量是防止模型崩溃的必要条件,且增长需超线性
  • 理论预测合成数据训练可能优于真实数据,适用于大规模模型研究

近年来,模型崩溃已成为语言模型训练中的关键问题,理解其内在机制至关重要。本文从概率视角研究递归参数化模型训练,旨在刻画模型崩溃发生的条件,并关键性地提出缓解方法。我们将递归训练过程建模为模型估计的随机游走,揭示样本量影响步长、估计方式决定游走方向与潜在偏差。在弱条件下,我们严格证明:在每一步训练中逐步增加样本量是防止模型崩溃的必要条件。特别地,当估计无偏时,所需增长速率为超线性;若存在显著估计偏差,该速率需进一步加快。基于此概率框架,我们还研究了在合成数据上递归训练获得优于仅用真实数据训练模型的概率。此外,我们将结果推广至渐近情形下的广义参数模型族。最后,通过大量模拟和真实数据集验证了理论结论。

原文摘要 · Abstract (English)

In recent years, model collapse has become a critical issue in language model training, making it essential to understand the underlying mechanisms driving this phenomenon. In this paper, we investigate recursive parametric model training from a probabilistic perspective, aiming to characterize the conditions under which model collapse occurs and, crucially, how it can be mitigated. We conceptualize the recursive training process as a random walk of the model estimate, highlighting how the sample size influences the step size and how the estimation procedure determines the direction and potential bias of the random walk. Under mild conditions, we rigorously show that progressively increasing the sample size at each training step is necessary to prevent model collapse. In particular, when the estimation is unbiased, the required growth rate follows a superlinear pattern. This rate needs to be accelerated even further in the presence of substantial estimation bias. Building on this probabilistic framework, we also investigate the probability that recursive training on synthetic data yields models that outperform those trained solely on real data. Moreover, we extend these results to general parametric model family in an asymptotic regime. Finally, we validate our theoretical results through extensive simulations and a real-world dataset.

模型崩溃概率建模语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。