arXiv:2506.07841cs.CVcs.AI2025-06被引 6

研究扩散模型在低噪声下的表现,揭示其生成与泛化的边界条件。

Diffusion models under low-noise regime

  • 在低噪声条件下分析扩散模型的去噪轨迹,发现不同训练集模型会偏离数据流形。
  • 训练集大小和数据几何结构显著影响模型去噪精度与得分估计能力。
  • 为实际应用中抗小扰动的生成模型可靠性提供新理解,适合关注模型稳健性的研究者。

近期研究表明,扩散模型在高噪声设置下表现出两种行为模式:记忆(再现训练数据)与泛化(生成新样本)。然而,当噪声水平较低时,扩散模型作为有效去噪器的行为仍不明确。为此,本文系统研究了低噪声扩散动态下的模型行为,探讨其对模型鲁棒性与可解释性的影响。基于(i)不同样本量的CelebA子集与(ii)解析的高斯混合基准,我们发现:即使高噪声输出趋于收敛,由不相交数据训练的模型在接近数据流形处仍出现显著发散。我们量化了训练集规模、数据几何结构及模型目标函数对去噪路径与得分准确率的影响,揭示了模型如何学习数据分布表示。该工作填补了对生成模型在现实场景中微小扰动下可靠性理解的空白。

原文摘要 · Abstract (English)

Recent work on diffusion models proposed that they operate in two regimes: memorization, in which models reproduce their training data, and generalization, in which they generate novel samples. While this has been tested in high-noise settings, the behavior of diffusion models as effective denoisers when the corruption level is small remains unclear. To address this gap, we systematically investigated the behavior of diffusion models under low-noise diffusion dynamics, with implications for model robustness and interpretability. Using (i) CelebA subsets of varying sample sizes and (ii) analytic Gaussian mixture benchmarks, we reveal that models trained on disjoint data diverge near the data manifold even when their high-noise outputs converge. We quantify how training set size, data geometry, and model objective choice shape denoising trajectories and affect score accuracy, providing insights into how these models actually learn representations of data distributions. This work starts to address gaps in our understanding of generative model reliability in practical applications where small perturbations are common.

扩散模型去噪生成模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。