发现扩散模型在迭代生成中从泛化转为记忆,提出用熵值选数据防退化。
A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- 从泛化到记忆的转变是模型退化的关键机制
- 合成数据熵值下降直接导致性能衰退
- 基于熵的数据筛选显著提升生成质量和多样性
扩散模型的广泛应用催生了大量生成数据,引发模型退化问题——即反复在合成数据上训练导致性能下降。以往研究多从方差缩小或分布偏移角度分析,但忽略了实际表现。本文揭示:在递归训练中,扩散模型会经历从泛化到记忆的转变,模型逐渐复制训练数据而非生成新内容。这一过程由每轮训练产生的合成数据熵值下降驱动,成为退化的明确指标。基于此,我们提出基于熵的数据选择策略,有效缓解从泛化到记忆的转变。实验表明,该方法显著提升递归生成的视觉质量与多样性,成功抑制模型退化。
原文摘要 · Abstract (English)
The widespread use of diffusion models has led to an abundance of AI-generated data, raising concerns about model collapse -- a phenomenon in which recursive iterations of training on synthetic data lead to performance degradation. Prior work primarily characterizes this collapse via variance shrinkage or distribution shift, but these perspectives miss practical manifestations of model collapse. This paper identifies a transition from generalization to memorization during model collapse in diffusion models, where models increasingly replicate training data instead of generating novel content during iterative training on synthetic samples. This transition is directly driven by the declining entropy of the synthetic training data produced in each training cycle, which serves as a clear indicator of model degradation. Motivated by this insight, we propose an entropy-based data selection strategy to mitigate the transition from generalization to memorization and alleviate model collapse. Empirical results show that our approach significantly enhances visual quality and diversity in recursive generation, effectively preventing collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。