arXiv:2503.12966cs.LGstat.ML2025-03JMLR被引 6

揭示了生成模型去噪策略的最优选择与数据规律的关系。

Optimal Denoising in Score-Based Generative Models: The Role of Data Regularity

  • 对比全去噪与半去噪策略,发现数据分布特性决定最优方法。
  • 规则数据下半去噪更优,奇异数据下全去噪性能更好。
  • 全去噪可缓解低维流形假设下的维度灾难问题。

基于得分的生成模型通过高斯噪声扰动分布后进行去噪实现顶尖采样性能。本文聚焦单一确定性去噪步骤,比较二次损失下的最优去噪器(称作'全去噪')与Hyvärinen(2025)提出的'半去噪'。研究表明,从分布距离角度评估性能时,不同数据假设导致截然不同的结论:对于足够规则的密度,半去噪优于全去噪;而对于奇异密度(如狄拉克混合或低维子空间支持的密度),全去噪更优。在后一种情况下,我们证明在线性流形假设下,全去噪可缓解维度灾难。

原文摘要 · Abstract (English)

Score-based generative models achieve state-of-the-art sampling performance by denoising a distribution perturbed by Gaussian noise. In this paper, we focus on a single deterministic denoising step, and compare the optimal denoiser for the quadratic loss, we name ''full-denoising'', to the alternative ''half-denoising'' introduced by Hyv{ä}rinen (2025). We show that looking at the performance in terms of distance between distributions tells a more nuanced story, with different assumptions on the data leading to very different conclusions. We prove that half-denoising is better than full-denoising for regular enough densities, while full-denoising is better for singular densities such as mixtures of Dirac measures or densities supported on a low-dimensional subspace. In the latter case, we prove that full-denoising can alleviate the curse of dimensionality under a linear manifold hypothesis.

生成模型去噪得分网络数据规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。