用联想记忆理论解释扩散模型从记忆到生成的转变。
Memorization to Generalization: Emergence of Diffusion Models from Associative Memory
- 将扩散模型视为记忆检索过程,用密集联想记忆框架分析其机制。
- 数据量增大时出现虚假状态,标志着模型从记忆转向泛化生成。
- 这些状态是生成能力的早期信号,适用于理解模型演化过程的研究者。
密集联想记忆(DenseAM)是霍普菲尔德网络的推广,具有更高的信息存储容量,可将训练数据点(记忆)存储在能量景观的局部极小值处。当训练数据量超过模型的临界存储容量时,会涌现出不同于训练数据的新局部极小值,称为‘虚假状态’,阻碍记忆恢复。本文从DenseAM视角审视扩散模型(DMs),将其生成过程视为一次记忆检索。在小数据情形下,DMs为每个训练样本创建独立吸引子,类似低于临界存储容量的DenseAM;随着数据量增加,模型由记忆转向泛化。我们识别出一个由DenseAM理论预测的临界中间相——虚假状态。在生成建模中,这些状态不再是负面产物,而是生成能力的最初迹象。我们刻画了这些被忽视状态的吸引盆地、能量景观曲率及计算特性,并在多种架构和数据集上验证其存在。
原文摘要 · Abstract (English)
Dense Associative Memories (DenseAMs) are generalizations of Hopfield networks, which have superior information storage capacity and can store training data points (memories) at local minima of the energy landscape. When the amount of training data exceeds the critical memory storage capacity of these models, new local minima, which are different from the training data, emerge. In Associative Memory these emergent local minima are called $\textit{spurious}\; \textit{states}$, which hinder memory retrieval. In this work, we examine diffusion models (DMs) through the DenseAM lens, viewing their generative process as an attempt of a memory retrieval. In the small data regimes, DMs create distinct attractors for each training sample, akin to DenseAMs below the critical memory storage. As the training data size increases, they transition from memorization to generalization. We identify a critical intermediate phase, predicted by DenseAM theory -- the spurious states. In generative modeling, these states are no longer negative artifacts but rather are the first signs of generative capabilities. We characterize the basins of attraction, energy landscape curvature, and computational properties of these previously overlooked states. Their existence is demonstrated across a wide range of architectures and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。