arXiv:2502.18865cs.LGcs.AI2025-02ICLR被引 15

揭示自消耗训练中模型崩溃的理论原因,给出稳定训练的条件。

A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops

  • 提出递归稳定性概念,分析架构与真实数据比例的影响。
  • 证明恒定比例的真实数据可确保Transformer收敛。
  • 为生成模型自训练提供理论指导,适合研究生成模型者。

高质量数据对训练大型生成模型至关重要,但在线真实数据已近乎枯竭。因此,模型越来越多地使用自身生成的数据进行后续训练,形成自消耗训练循环(STL)。然而,实证结果差异显著:部分模型退化甚至崩溃,而另一些则成功避免失败,理论理解存在明显空白。本文引入递归稳定性概念,首次建立理论泛化分析,揭示模型架构及真实与合成数据比例如何影响STL的成功。进一步将分析扩展至上下文学习中的Transformer,证明即使真实数据占比恒定也能实现收敛,并给出合成数据最优规模的启示。

原文摘要 · Abstract (English)

High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of recursive stability and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing.

生成模型自训练理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。