揭示扩散模型生成样本的几何演化规律,解释其泛化机制。
Diffusion Model's Generalization Can Be Characterized by Inductive Biases toward a Data-Dependent Ridge Manifold
- 构建数据依赖的平滑流形,刻画生成过程的几何路径。
- 生成样本按进入、对齐、沿流形滑动三阶段演化,误差分量决定轨迹。
- 在合成数据与MNIST上验证,适用于低维与高维场景。
我们研究一种数据依赖的扩散模型泛化概念:当模型不记忆训练集时,其生成样本相对于数据诱导几何结构的位置如何?为此,我们引入由平滑经验分布构造的时间依赖型对数密度脊流形,并用于刻画逆时间推断。主要结果表明,生成样本经历‘进入-对齐-滑动’机制:先进入脊流形邻域,其到流形的距离由训练误差的法向分量控制,沿流形的运动则由切向分量主导。我们进一步通过误差的方向分解将该几何图像与训练动态关联,对随机特征模型明确建立了架构偏差与优化误差的定量分离。在合成多模态数据和MNIST潜在扩散模型上的实验支持了预测的几何行为,涵盖低维与高维情形。
原文摘要 · Abstract (English)
We study a data-dependent notion of diffusion-model generalization: when a model does not memorize the training set, where do its generated samples go relative to the geometry induced by the data? To answer this, we introduce a time-dependent family of log-density ridge manifolds constructed from the smoothed empirical distribution, and use it to characterize reverse-time inference. Our main result shows that generated samples evolve by a reach-align-slide mechanism: they first enter a neighborhood of the ridge, then their distance to the ridge is controlled by the normal component of training error, and finally their motion along the ridge is controlled by the tangential component. We further connect this geometric picture to training dynamics through directional decompositions of the learned error, and make this link explicit for random feature models, where architectural bias and optimization error can be separated quantitatively. Experiments on synthetic multimodal data and MNIST latent diffusion support the predicted geometric behavior in both low and high dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。