首次给出扩散模型的算法与数据相关泛化界,揭示优化过程对生成质量的关键影响。
Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models
- 从优化动态出发,建立依赖训练算法和数据分布的泛化边界
- 实验表明优化超参数显著影响生成分布的泛化能力
- 为理解扩散模型实际表现提供新理论视角,适合研究生成模型泛化者
基于分数的生成模型(SGMs)已成为最受欢迎的生成模型之一。现有大量研究关注SGM的离散化或统计性能分析,通常在不同度量下推导真实数据分布与模型生成分布之间的界限,常显示随训练样本数呈多项式收敛。然而这些方法多采用近似理论视角,往往过于悲观且粗糙,难以解释SGM的实际成功,也未充分考虑实践中用于训练分数网络的优化算法的作用。为此,我们首先通过简单实验展示优化超参数对生成分布泛化能力的具体影响。本文旨在填补这一理论空白,首次为SGMs提供算法与数据相关的泛化分析。具体而言,我们建立了显式包含学习算法优化动态的边界,为理解SGM的泛化行为提供了新见解。理论结果在多个数据集上得到实证支持。
原文摘要 · Abstract (English)
Score-based generative models (SGMs) have emerged as one of the most popular classes of generative models. A substantial body of work now exists on the analysis of SGMs, focusing either on discretization aspects or on their statistical performance. In the latter case, bounds have been derived, under various metrics, between the true data distribution and the distribution induced by the SGM, often demonstrating polynomial convergence rates with respect to the number of training samples. However, these approaches adopt a largely approximation theory viewpoint, which tends to be overly pessimistic and relatively coarse. In particular, they fail to fully explain the empirical success of SGMs or capture the role of the optimization algorithm used in practice to train the score network. To support this observation, we first present simple experiments illustrating the concrete impact of optimization hyperparameters on the generalization ability of the generated distribution. Then, this paper aims to bridge this theoretical gap by providing the first algorithmic- and data-dependent generalization analysis for SGMs. In particular, we establish bounds that explicitly account for the optimization dynamics of the learning algorithm, offering new insights into the generalization behavior of SGMs. Our theoretical findings are supported by empirical results on several datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。