arXiv:2602.16601stat.MLcs.LG2026-02中稿 · ICML被引 3

揭示扩散模型在合成数据迭代训练中的误差累积规律

Quantifying Error Propagation and Model Collapse in Diffusion Models

  • 从理论推导出生成分布与目标分布的偏差上下界
  • 发现多轮训练后偏差呈误差折扣求和,受新数据比例影响
  • 首次给出标准扩散模型的下界,适合研究模型退化者

机器学习模型越来越多地在合成数据上训练或微调。递归使用此类数据已被观察到在多种任务中显著降低性能,通常表现为生成分布逐步偏离目标分布。本文针对基于得分的扩散模型,理论上分析了这一现象。对于一个每轮训练均结合合成数据与来自目标分布的新样本的真实流程,我们给出了生成分布与目标分布间累积偏移的上下界。值得注意的是,据我们所知,这是首个针对标准扩散模型的生成分布与目标分布之间偏差的下界。我们的结果可刻画不同漂移阶段,取决于得分估计误差及每轮使用的新增样本比例。在某一特定阶段,经过多轮再训练后的累积偏差可表示为各轮得分估计误差的折扣求和。我们在合成数据和图像上提供了实证结果以验证理论。

原文摘要 · Abstract (English)

Machine learning models are increasingly trained or fine-tuned on synthetic data. Recursively training on such data has been observed to significantly degrade performance in a wide range of tasks, often characterized by a progressive drift away from the target distribution. In this work, we theoretically analyze this phenomenon in the setting of score-based diffusion models. For a realistic pipeline where each training round uses a combination of synthetic data and fresh samples from the target distribution, we obtain upper and lower bounds on the accumulated divergence between the generated and target distributions. Notably, to the best of our knowledge, this is the first lower bound on the divergence between the learned and target distributions, even for standard diffusion models. Our results allow us to characterize different regimes of drift, depending on the score estimation error and the proportion of fresh data used in each generation. In a certain regime, the accumulated divergence after several retraining rounds can be expressed as a discounted sum of score estimation errors made at each generation. We also provide empirical results on synthetic data and images to illustrate the theory.

扩散模型误差传播模型退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。