arXiv:2605.19483cs.LG2026-05

从动态系统视角解释生成模型训练中的记忆现象

Adynamical systems view of training generativemodels and the memorization phenomenon

  • 基于双时间尺度随机梯度下降的动态特性分析
  • 揭示生成模型在训练中出现长时间输出重复的机制
  • 适合关注模型训练动态与过拟合现象的研究者

结合作者之一VSB关于生成模型坍塌现象及高维随机梯度下降中双时间尺度动力学的最新研究,本文从系统理论角度解释生成模型中的记忆现象。该分析完全基于训练过程的动态特性。具体而言,我们引用Austin [2016] 的结果,提出一种损失函数的简化模型:其对部分变量有强依赖,对其他变量依赖较弱。这自然导致在常步长SGD中存在两个明显不同的时间尺度,这一现象此前已被用于解释SGD中的双下降现象(Borkar [2026])。结合Borkar [2025a] 提出的SGD坍塌数学模型,以及Azizian等[2024]的最新成果,我们分析了常步长SGD,解释了生成模型在持续调参过程中产生相同或相似输出、长时间维持不变的现象。该研究为机器学习文献中多个现象提供了基于动力系统的新视角及其内在联系。

原文摘要 · Abstract (English)

Using recent works of one of the authors (VSB) on collapse in generative models and two time scale dynamics in stochastic gradient descent in high dimensions, we give a system theoretic explanation of the memorization phenomenon in generative models. This relies purely on the dynamic aspects of the training phase. Specifically, we use a result of Austin [2016] to motivate a stylized model for the loss function for stochastic gradient descent (SGD) wherein the loss function has a strong dependence on some variables and weak dependence on the rest in a precise sense. This naturally leads to two distinct time scales in the constant step size SGD that is commonly used in machine learning. This fact has been used to explain the double descent phenomenon in SGD in Borkar [2026]. In conjunction with a mathematical model for collapse phenomenon in SGD developed in Borkar [2025a], we analyze the constant step size SGD using the recent results of Azizian et al. [2024] in order to explain the phenomenon of memorization wherein a generative model that is concurrently being tuned yields the same or similar outputs for significant stretches of time. This gives a novel perspective on the aforementioned phenomena reported in machine learning literature and their interrelationships, using a dynamical systems viewpoint.

生成模型记忆现象动态系统SGD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。