arXiv:2508.14871cs.LGcs.CV2025-08

通过数据感知的噪声调节,提升扩散模型生成质量。

Squeezed Diffusion Models

  • 基于主成分方向异步缩放噪声,实现数据驱动的噪声注入。
  • 在CIFAR-10/100和CelebA-64上FID降低最高15%,召回率提升。
  • 无需修改架构,简单噪声设计即可显著改善生成效果。

扩散模型通常注入各向同性的高斯噪声,忽略了数据结构。受量子压缩态按海森堡不确定性原理重分配不确定性的启发,我们提出压缩扩散模型(SDM),沿训练分布主成分方向非均匀缩放噪声。由于压缩在物理中可提升信噪比,我们假设以数据依赖方式缩放噪声能更好辅助扩散模型学习关键特征。研究了两种配置:(i) 海森堡扩散模型,在主轴上缩放的同时对正交方向反向缩放;(ii) 标准SDM变体,仅在主轴上缩放。出人意料的是,在CIFAR-10/100和CelebA-64上,轻微反压缩(即主轴方差增大)持续提升FID最高达15%,并将精确率-召回率前沿推向更高召回率。结果表明,简单的数据感知噪声调控可在不改变架构的前提下带来稳健的生成性能提升。

原文摘要 · Abstract (English)

Diffusion models typically inject isotropic Gaussian noise, disregarding structure in the data. Motivated by the way quantum squeezed states redistribute uncertainty according to the Heisenberg uncertainty principle, we introduce Squeezed Diffusion Models (SDM), which scale noise anisotropically along the principal component of the training distribution. As squeezing enhances the signal-to-noise ratio in physics, we hypothesize that scaling noise in a data-dependent manner can better assist diffusion models in learning important data features. We study two configurations: (i) a Heisenberg diffusion model that compensates the scaling on the principal axis with inverse scaling on orthogonal directions and (ii) a standard SDM variant that scales only the principal axis. Counterintuitively, on CIFAR-10/100 and CelebA-64, mild antisqueezing - i.e. increasing variance on the principal axis - consistently improves FID by up to 15% and shifts the precision-recall frontier toward higher recall. Our results demonstrate that simple, data-aware noise shaping can deliver robust generative gains without architectural changes.

扩散模型噪声调度生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。