在弱假设下证明了分数模型可逼近任意亚高斯分布,性能接近最优。
Approximation and Generalization Abilities of Score-based Neural Network Generative Models for Sub-Gaussian Distributions
- 仅假设分布为α-亚高斯,无需平滑或密度下界条件
- 深度神经网络以n^{-1}误差率逼近得分函数,达到近最优收敛速度
- 适用于低维到高维场景,适合对理论性能有要求的研究者
本文研究了在d维空间中从n个独立同分布观测值估计未知分布$P_0$时,基于分数的神经网络生成模型(SGMs)的逼近与泛化能力。仅假设$P_0$为α-亚高斯分布,在任意时间步$t \in [t_0, n^{\mathcal{O}(1)}]$下,其中$t_0 > \mathcal{O}(α^2n^{-2/d}\log n)$,存在一个宽度不超过$\mathcal{O}(n^{3/d}\log_2n)$、深度不超过$\mathcal{O}(\log^2n)$的深层ReLU神经网络,能以$\tilde{\mathcal{O}}(n^{-1})$的均方误差逼近得分函数,并在分数匹配损失下实现近最优的$\tilde{\mathcal{O}}(n^{-1}t_0^{-d/2})$收敛速率。该框架具有普适性,可在比以往工作更弱的条件下建立收敛速率。若进一步假设目标密度$ p_0 $属于Sobolev或Besov类,配合适当的早停策略,神经网络型SGMs可达到几乎最小最大收敛速率(对数因子内)。分析去除了若干关键假设,如得分函数的Lipschitz连续性或目标密度的严格正下界。
原文摘要 · Abstract (English)
This paper studies the approximation and generalization abilities of score-based neural network generative models (SGMs) in estimating an unknown distribution $P_0$ from $n$ i.i.d. observations in $d$ dimensions. Assuming merely that $P_0$ is $α$-sub-Gaussian, we prove that for any time step $t \in [t_0, n^{\mathcal{O}(1)}]$, where $t_0 > \mathcal{O}(α^2n^{-2/d}\log n)$, there exists a deep ReLU neural network with width $\leq \mathcal{O}(n^{\frac{3}{d}}\log_2n)$ and depth $\leq \mathcal{O}(\log^2n)$ that can approximate the scores with $\tilde{\mathcal{O}}(n^{-1})$ mean square error and achieve a nearly optimal rate of $\tilde{\mathcal{O}}(n^{-1}t_0^{-d/2})$ for score estimation, as measured by the score matching loss. Our framework is universal and can be used to establish convergence rates for SGMs under milder assumptions than previous work. For example, assuming further that the target density function $p_0$ lies in Sobolev or Besov classes, with an appropriately early stopping strategy, we demonstrate that neural network-based SGMs can attain nearly minimax convergence rates up to logarithmic factors. Our analysis removes several crucial assumptions, such as Lipschitz continuity of the score function or a strictly positive lower bound on the target density.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。