arXiv:2606.19876cs.LGmath.OC2026-06

改用反向费雪散度可让梯度下降全局收敛,突破传统方法的初始化依赖问题。

Global Convergence of Gradient Descent for Score Matching in Gaussian Mixtures via Reverse Fisher Divergence

  • 采用反向费雪散度作为目标函数,使优化过程更稳定
  • 在单高斯教师与混合高斯学生模型下,任意初始化均全局收敛
  • 适用于需要稳定训练的生成模型与无归一化建模任务

得分匹配是现代生成建模、扩散模型、无归一化统计模型和逆问题的核心训练目标。传统方法最小化前向费雪散度(期望基于教师分布),但在简单高斯混合模型中仍存在不良且依赖初始化的收敛行为。本文研究一种替代目标:反向费雪散度(期望基于学生分布)。分析梯度下降在拟合高斯混合模型时的表现,证明该改变显著改善优化性质。当教师为单高斯而学生为固定权重、单位协方差的高斯混合模型时,可从任意初始点实现全局收敛。进一步扩展至教师也为高斯混合模型的情形,在全局随机初始化及目标均值间$ ildeigOmega(1)$分离假设下,证明全局收敛性;高概率下每个学生成分收敛至最近的教师成分,并给出学生分布以总变差距离收敛的条件。证明依赖于新的基于李雅普诺夫函数的梯度下降动力学分析,表明反向费雪散度具有比前向更优的优化景观。

原文摘要 · Abstract (English)

The score matching problem is a central training objective in modern generative modeling, diffusion models, fitting unnormalized statistical models, and inverse problems. A standard approach is to minimize the forward Fisher divergence, where the expectation is taken with respect to the teacher distribution. However, recent results show that even in simple Gaussian mixture model settings, this objective can lead to undesirable and initialization-dependent convergence behavior. In this paper, we study an alternative objective: the reverse Fisher divergence, where the expectation is taken with respect to the student distribution. We analyze gradient descent (GD) for fitting Gaussian mixture models and show that this change in the objective leads to significantly better optimization properties. First, when the teacher distribution is a single Gaussian and the student is a Gaussian mixture model with fixed weights and identity covariances, we prove the global convergence of GD from arbitrary initializations. Second, we extend the analysis to the case where the teacher is also a Gaussian mixture model and prove global convergence guarantees under a global random initialization scheme and a $\widetildeΩ(1)$-separation assumption on the target means. In particular, with high probability, each student component converges near its closest teacher component, and we provide conditions under which the student distribution converges in total variation distance. Our proofs rely on a new Lyapunov-based analysis of the gradient descent dynamics, showing that the reverse Fisher divergence has a much more favorable optimization landscape than the forward Fisher divergence.

生成模型梯度下降高斯混合优化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。