AI模型自训练时如何避免数据污染?理论证明它能收敛到真实分布。
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
- 在温和条件下,模型可收敛至真实数据分布,不受具体模型结构影响。
- 收敛速度取决于模型内在速率与每轮真实数据占比的较小值,存在数据与模型限制的相变点。
- 修正真实数据偏差可防止偏差在训练中被放大,适合长期稳定AI系统设计者参考。
随着人工智能生成内容泛滥,模型越来越多地使用自身输出进行训练,可能导致性能逐步退化甚至崩溃。本文首次在理论上证明,在模型无关的温和条件下,模型仍可收敛至真实数据生成分布。收敛速率由模型内在速率与每轮训练中真实数据占比的最小值决定,揭示了数据受限与模型受限之间的相变现象。进一步表明,若真实数据存在偏差,通过纠正偏差可有效阻止早期偏差在训练过程中持续积累和放大。大量模拟、真实图像与文本实验验证了该理论框架,为复杂污染环境下人工智能系统的长期稳定性提供了定量条件。
原文摘要 · Abstract (English)
As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse. In this article, we provide the first positive, rigorous theoretical results, to the best of our knowledge, showing that under model-agnostic mild conditions, the model converges to the true data-generating distribution. The convergence rate is the minimum of the model's intrinsic rate and the fraction of real data at each training iteration, revealing a phase transition between data-limited and model-limited regimes. We further show that, for biased real data, correcting the bias prevents the persistence and amplification of early bias over training iteration. Extensive experiments across simulations, real images and texts validate our theoretical framework, establishing quantitative conditions for long-term AI stability in contaminated environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。