揭示深度生成模型在重尾分布上的能力局限,打破其万能生成的误解。
On the Statistical Capacity of Deep Generative Models
- 通过测度集中与凸几何工具,证明常见生成模型无法生成重尾样本
- 在高斯潜变量下,模型只能生成轻尾集中样本,误差不可无限逼近零
- 结果适用于变分自编码器、生成对抗网络及扩散模型,适合关注生成边界的研究者
深度生成模型广泛用于从复杂高维分布中采样,尽管表现成功,其统计性质仍不清晰。普遍假设认为,只要训练数据充足且网络足够大,生成样本可任意接近任意连续目标分布。本文建立统一框架,推翻该信念:包括变分自编码器和生成对抗网络在内的多种模型,并非万能生成器。在主流高斯潜变量假设下,这些模型仅能生成集中且轻尾的样本。借助测度集中与凸几何工具,我们对更一般的对数凹与强对数凹潜变量分布得出了类似结论。通过归约论证,将结果扩展至扩散模型。利用Gromov-Levy不等式,当潜变量位于正里奇曲率流形时也获得相似保证。这些结果揭示了常见生成模型在处理重尾分布上的能力瓶颈。通过模拟与金融数据验证了理论的实证相关性。
原文摘要 · Abstract (English)
Deep generative models are routinely used in generating samples from complex, high-dimensional distributions. Despite their apparent successes, their statistical properties are not well understood. A common assumption is that with enough training data and sufficiently large neural networks, deep generative model samples will have arbitrarily small errors in sampling from any continuous target distribution. We set up a unifying framework that debunks this belief. We demonstrate that broad classes of deep generative models, including variational autoencoders and generative adversarial networks, are not universal generators. Under the predominant case of Gaussian latent variables, these models can only generate concentrated samples that exhibit light tails. Using tools from concentration of measure and convex geometry, we give analogous results for more general log-concave and strongly log-concave latent variable distributions. We extend our results to diffusion models via a reduction argument. We use the Gromov--Levy inequality to give similar guarantees when the latent variables lie on manifolds with positive Ricci curvature. These results shed light on the limited capacity of common deep generative models to handle heavy tails. We illustrate the empirical relevance of our work with simulations and financial data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。