arXiv:2410.08727stat.MLcs.LG2024-10被引 27

扩散模型在数据稀少时会逐步记忆,先记关键特征,再记细节。

Losing dimensions: Geometric memorization in generative diffusion

  • 通过学习得分场测量潜在维度,发现模型逐渐失去变化能力。
  • 数据越少,模型越集中在少数样本上,其他变化被冻结。
  • 揭示了从泛化到复制的几何记忆新阶段,适合研究生成模型机理者阅读。

扩散模型驱动主流生成AI,但其在低维流形上何时及如何记忆训练数据仍不明确。我们发现记忆是渐进出现而非突然发生:当数据稀缺时,扩散模型经历平滑坍缩,其在独立方向上的表达能力逐渐减弱。通过学习得分场测量潜在维度,我们揭示生成行为逐渐集中于少数样本,其他变化逐渐“冻结”。我们提出几何记忆理论,表明显著特征首先坍缩,随后是细粒度细节,最终导致近乎点对点的复制。这一过程类似于物理系统凝聚为少数低能态。理论预测在合成与真实数据上均成立,将几何记忆识别为介于泛化与精确复制之间的独特相变阶段。

原文摘要 · Abstract (English)

Diffusion models power leading generative AI, but when and how they memorize training data, especially on low-dimensional manifolds, remains unclear. We find memorization emerges gradually, not abruptly: as data become scarce, diffusion models experience a smooth collapse where their capacity to vary across independent directions diminishes. Measuring latent dimensionality via the learned score field, we reveal how generative behavior increasingly centers on a few examples while other variations "freeze out". We propose a geometric memorization theory, showing that salient features collapse first, then finer details, leading to near point-wise replication. This mirrors physical systems condensing into a few low-energy configurations. Our theoretical predictions align with both synthetic and real data, identifying geometric memorization as a distinct phase between generalization and exact copying.

扩散模型生成模型几何记忆数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。