研究生成模型在迭代训练中为何不会崩溃,揭示避免崩溃的条件。
When Models Don't Collapse: On the Consistency of Iterative MLE
- 在真实数据渐增合成数据的设定下,证明模型可避免崩溃。
- 即使真实数据占比趋近于零,非渐近界仍保证性能不降。
- 首次严格证明迭代生成模型在特定条件下会极速崩溃。
生成模型的广泛应用形成了反馈循环:每一代模型都基于前代生成的部分合成数据进行训练。这引发了对模型崩溃的担忧——即因反复训练于合成数据而导致性能严重退化。然而,现有文献对此现象的严重性存在不同结论,其影响范围与规避条件尚不明确。为此,本文在标准假设下(类似传统MLE渐近一致性和正态性的假设),对最大似然估计(MLE)中的模型崩溃进行理论分析。研究发现,在合成数据逐步加入原始数据集的自然设定中,即使真实数据比例趋于零,仍可通过非渐近界避免崩溃。另一方面,证明了若超出MLE一致性之外的某些假设不成立,则模型崩溃可能以任意速度发生,即便原始数据仍存在于训练集中。据我们所知,这是首个关于累积数据下迭代生成建模导致快速崩溃的严格示例。
原文摘要 · Abstract (English)
The widespread use of generative models has created a feedback loop, in which each generation of models is trained on data partially produced by its predecessors. This process has raised concerns about model collapse: A critical degradation in performance caused by repeated training on synthetic data. However, different analyses in the literature have reached different conclusions as to the severity of model collapse. As such, it remains unclear how concerning this phenomenon is, and under which assumptions it can be avoided. To address this, we theoretically study model collapse for maximum likelihood estimation (MLE), in a natural setting where synthetic data is gradually added to the original data set. Under standard assumptions (similar to those long used for proving asymptotic consistency and normality of MLE), we establish non-asymptotic bounds showing that collapse can be avoided even as the fraction of real data vanishes. On the other hand, we prove that some assumptions (beyond MLE consistency) are indeed necessary: Without them, model collapse can occur arbitrarily quickly, even when the original data is still present in the training set. To the best of our knowledge, these are the first rigorous examples of iterative generative modeling with accumulating data that rapidly leads to model collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。