分析过参数化模型学习单高斯分布的梯度下降收敛性,揭示初始化与噪声水平的影响。
Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian
- 用过参数化学生模型在真实高斯数据上训练,研究梯度下降动态
- 大噪声下全局收敛;小噪声下需指数级小初始化才收敛
- 随机初始化时仅一个参数收敛但损失仍以1/τ速率下降
得分匹配已成为现代生成建模的核心训练目标,尤其在扩散模型中用于通过估计得分函数学习高维数据分布。尽管其在实践中取得成功,但对得分匹配优化行为的理论理解,特别是在过参数化情形下的表现仍有限。本文研究过参数化模型在单个真实高斯分布数据上使用总体得分匹配目标进行梯度下降训练的动力学。当噪声尺度足够大时,证明了梯度下降的全局收敛性;在低噪声情况下,识别出存在不动点,表明全局收敛难以证明。然而,当参数被指数级小初始化时,梯度下降可确保所有参数收敛至真实值。若无此初始化,则参数可能不收敛。此外,当参数从远离真实值的高斯分布随机初始化时,以高概率仅一个参数收敛,其余发散,但损失仍以1/τ的速率收敛至零,其中τ为迭代次数。我们还建立了该情形下几乎匹配的收敛速率下界。这是首个在得分匹配框架下对至少三个成分高斯混合实现全局收敛保证的工作。
原文摘要 · Abstract (English)
Score matching has become a central training objective in modern generative modeling, particularly in diffusion models, where it is used to learn high-dimensional data distributions through the estimation of score functions. Despite its empirical success, the theoretical understanding of the optimization behavior of score matching, particularly in over-parameterized regimes, remains limited. In this work, we study gradient descent for training over-parameterized models to learn a single Gaussian distribution. Specifically, we use a student model with $n$ learnable parameters and train it on data generated from a single ground-truth Gaussian using the population score matching objective. We analyze the optimization dynamics under multiple regimes. When the noise scale is sufficiently large, we prove a global convergence result for gradient descent. In the low-noise regime, we identify the existence of a stationary point, highlighting the difficulty of proving global convergence in this case. Nevertheless, we show convergence under certain initialization conditions: when the parameters are initialized to be exponentially small, gradient descent ensures convergence of all parameters to the ground truth. We further prove that without the exponentially small initialization, the parameters may not converge to the ground truth. Finally, we consider the case where parameters are randomly initialized from a Gaussian distribution far from the ground truth. We prove that, with high probability, only one parameter converges while the others diverge, yet the loss still converges to zero with a $1/τ$ rate, where $τ$ is the number of iterations. We also establish a nearly matching lower bound on the convergence rate in this regime. This is the first work to establish global convergence guarantees for Gaussian mixtures with at least three components under the score matching framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。