arXiv:2605.21402stat.MLcond-mat.dis-nn2026-05被引 2

揭示生成模型从记忆到泛化的关键转变机制

Memorisation, convergence and generalisation in generative models

论文配图:Memorisation, convergence and generalisation in generative models
图 1 · 摘自论文原文
  • 通过线性生成模型分析数据量与收敛关系
  • 当样本数线性于输入维度时,模型开始收敛但不恢复主成分
  • 泛化包含两个独立目标:匹配分布整体和恢复核心因子

生成神经网络从有限但大量的样本中学习生成高度逼真的图像——它们是真正学习了数据分布,还是仅仅记忆了训练集?为解答此问题,Kadkhodaie、Guth、Simoncelli 和 Mallat(ICLR '24)在数据集的不同子集上独立训练扩散模型,发现当训练样本足够多时,模型收敛到几乎相同的密度。这引出两个基本问题:达到收敛需要多少数据?收敛反映了对数据分布学习的哪些方面?本文通过线性生成模型提供该转变过程的精确解析表征,发现模型在低负载时会记忆数据,而当样本数与输入维度成线性关系时,收敛现象连续出现。令人惊讶的是,收敛对恢复数据主潜因子不敏感,后者呈现尖锐过渡。将方法扩展至幂律谱数据后,在卷积去噪器实验及Kadkhodaie等人的数据中均观察到相同区分:收敛与潜因子恢复属于不同机制。因此,生成模型的泛化可分解为至少两个独立目标:拟合数据分布主体,以及恢复主潜因子。这两个目标对应于真实与学习分布之间的不同距离度量,且仅前者被收敛所捕捉。

原文摘要 · Abstract (English)

Generative neural networks learn how to produce highly realistic images from a large, but finite number of examples - or do they simply memorise their training set? To settle this question, Kadkhodaie, Guth, Simoncelli and Mallat (ICLR '24) trained diffusion models independently on disjoint subsets of a dataset and showed that they converge to nearly the same density when the number of training images is large enough. This result raises two basic questions: how much data do you need for convergence, and what does convergence capture about learning the data distribution? Here, we address these questions by providing an exact analytical characterisation of the transition from memorisation to generalisation in linear generative models. We find that these models memorise at small load, while convergence emerges continuously when the number of samples is linear in the input dimension. Strikingly, we find that convergence is insensitive to recovery of the principal latent factors of the data, which are recovered in a sharp transition. After extending our approach to data with power-law spectra, we find the same distinction between convergence and latent recovery in our experiments with convolutional denoisers and in the data of Kadkhodaie et al. We thus show that generalisation in generative models decomposes into at least two distinct objectives: matching the bulk of the data distribution and recovering the principal latent factors. These objectives correspond to two different distances between true and learnt data distribution, and only the first one is captured by convergence.

生成模型泛化能力分布学习潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。