揭示扩散模型隐式正则化的算法根源,解析为何能泛化。
Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis
- 从算法稳定性出发,提出得分稳定性概念衡量模型对数据扰动的敏感度。
- 发现早停、粗粒度采样器离散和SGD优化均带来隐式正则化作用。
- 突破传统结构依赖分析,为扩散模型泛化提供新理论视角,适合研究者参考。
扩散模型的成功引发了对其高维设置下泛化行为的重要疑问。已有研究表明,若训练与采样完全理想,模型会记忆训练数据,表明正则化对泛化至关重要。现有理论分析多依赖于算法无关的统一收敛技术,且过度依赖模型结构来获得泛化界。本文转而关注促进泛化的算法特性,构建了扩散模型的算法依赖泛化理论。借鉴算法稳定性框架,提出得分稳定性概念,量化得分匹配算法对数据扰动的敏感性。我们基于得分稳定性推导出泛化界,并应用于多个基础学习场景,识别出多种正则化来源。特别地,考虑了带早停的去噪得分匹配(去噪正则化)、采样器范围内的粗粒度离散化(采样器正则化)以及使用SGD优化(优化正则化)。通过将分析建立在算法属性而非模型结构之上,我们识别出扩散模型中此前文献未注意到的多种隐式正则化机制。
原文摘要 · Abstract (English)
The success of denoising diffusion models raises important questions regarding their generalisation behaviour, particularly in high-dimensional settings. Notably, it has been shown that when training and sampling are performed perfectly, these models memorise training data -- implying that some form of regularisation is essential for generalisation. Existing theoretical analyses primarily rely on algorithm-independent techniques such as uniform convergence, heavily utilising model structure to obtain generalisation bounds. In this work, we instead leverage the algorithmic aspects that promote generalisation in diffusion models, developing a general theory of algorithm-dependent generalisation for this setting. Borrowing from the framework of algorithmic stability, we introduce the notion of score stability, which quantifies the sensitivity of score-matching algorithms to dataset perturbations. We derive generalisation bounds in terms of score stability, and apply our framework to several fundamental learning settings, identifying sources of regularisation. In particular, we consider denoising score matching with early stopping (denoising regularisation), sampler-wide coarse discretisation (sampler regularisation) and optimising with SGD (optimisation regularisation). By grounding our analysis in algorithmic properties rather than model structure, we identify multiple sources of implicit regularisation unique to diffusion models that have so far been overlooked in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。