反复用自生成数据训练模型会引发崩溃,这是统计规律而非偶然。
A Note on Shumailov et al. (2024): `AI Models Collapse When Trained on Recursively Generated Data'
- 用核密度估计拟合数据并反复采样,模拟模型迭代训练过程。
- 实验表明模型崩溃是不可避免的统计现象,与具体模型无关。
- 适合关注大模型训练稳定性和数据循环风险的研究者阅读。
Shumailov等人(2024)的研究表明,反复用合成数据训练生成模型会导致模型崩溃。这一发现引发了广泛关注与讨论,尤其在当前模型几乎耗尽真实数据的背景下。本文通过拟合数据分布(采用核密度估计,KDE)并重复采样,探究该现象背后的机制。结果表明,所观察到的崩溃现象本质上是统计学上的必然结果,可能无法避免。
原文摘要 · Abstract (English)
The study conducted by Shumailov et al. (2024) demonstrates that repeatedly training a generative model on synthetic data leads to model collapse. This finding has generated considerable interest and debate, particularly given that current models have nearly exhausted the available data. In this work, we investigate the effects of fitting a distribution (through Kernel Density Estimation, or KDE) or a model to the data, followed by repeated sampling from it. Our objective is to develop a theoretical understanding of the phenomenon observed by Shumailov et al. (2024). Our results indicate that the outcomes reported are a statistical phenomenon and may be unavoidable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。