解析生成模型训练中随机梯度下降的收敛性,给出理论指导。
Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models
- 针对一般参数化,给出带权重的得分匹配目标下SGD的非凸收敛率。
- 对两层ReLU网络,通过神经正切核分析得到训练过程中的得分逼近误差。
- 揭示重加权因子对误差的影响,为实际训练策略提供理论依据。
基于得分的生成模型(SGMs)在数据生成任务中表现出色。尽管其采样过程的统计性质日益清晰,但训练背后的优化动态仍不明确。SGMs通常通过最小化加权去噪得分匹配目标进行训练,但使用随机梯度时的优化保证仍有限。本文研究了SGD在SGMs中的应用,在两种互补情形下取得成果:首先,针对一般得分参数化,建立了加权去噪得分匹配目标上SGD的非凸收敛速率,明确包含依赖于训练调度的权重因子;其次,针对过参数化的两层ReLU网络,构建了适用于扩散训练与随机梯度的神经正切核分析,获得沿SGD轨迹的得分近似误差界。最后,我们的分析量化了重加权因子对得分近似误差的作用,为实践中权重选择提供了理论指导。
原文摘要 · Abstract (English)
Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising scorematching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. First, for general score parameterizations, we establish a non-convex convergence rate for SGD on the weighted denoising score-matching objective, with explicit dependence on the schedule-dependent weighting factors. Second, for overparameterized two-layer ReLU networks, we develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Finally, our analysis quantifies the role of the reweighting factor in the score approximation error, providing theoretical guidance for weighting choices used in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。