arXiv:2512.11867cs.LGcs.AI2025-12被引 1

重复使用生成数据训练会引发分布漂移,导致模型性能下降。

On the Dangers of Bootstrapping Generation for Continual Learning and Beyond

  • 用生成数据反复训练,引入显著偏差与方差。
  • 主流生成模型在重复训练中出现崩溃,隐空间对齐失效。
  • 警示持续学习中滥用合成数据的风险,适合研究者参考。

合成数据用于模型训练日益普遍。尽管可扩充数据,但反复使用生成数据会因数据污染引发分布漂移与性能退化。本文从持续学习视角考察此自举过程,关联生成经验回放(GER)方法。通过统计分析表明,合成数据使训练目标引入显著偏差与方差,削弱最大似然估计可靠性。实证显示,主流生成模型在重复训练下发生崩溃。量化结果表明,当前最先进的GER方法无法保持隐空间对齐。研究揭示了合成数据在持续学习中的严重风险。

原文摘要 · Abstract (English)

The use of synthetically generated data for training models is becoming a common practice. While generated data can augment the training data, repeated training on synthetic data raises concerns about distribution drift and degradation of performance due to contamination of the dataset. We investigate the consequences of this bootstrapping process through the lens of continual learning, drawing a connection to Generative Experience Replay (GER) methods. We present a statistical analysis showing that synthetic data introduces significant bias and variance into training objectives, weakening the reliability of maximum likelihood estimation. We provide empirical evidence showing that popular generative models collapse under repeated training with synthetic data. We quantify this degradation and show that state-of-the-art GER methods fail to maintain alignment in the latent space. Our findings raise critical concerns about the use of synthetic data in continual learning.

持续学习生成数据分布漂移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。