arXiv:2605.16415cs.CVcs.LG2026-05

扩散模型的创造力源于去噪器架构与数据分布的相互作用。

Diffusion Models, Denoiser Architecture and Creativity

论文配图:Diffusion Models, Denoiser Architecture and Creativity
图 1 · 摘自论文原文
  • 分析三种去噪器架构的生成分布,揭示其与目标分布的关系。
  • 微调UNet架构会显著改变生成结果的现实感与创意性。
  • 模型成功依赖于去噪器先验与真实数据分布的高度契合。

扩散模型的创造力指其生成与训练数据不同的高度逼真图像的能力。这一现象看似矛盾,因为若去噪器是给定训练集的贝叶斯最优去噪器,则模型将仅复制训练样本。本文通过理论与实证研究发现,创造力源于去噪器架构与目标分布之间的交互作用。理论上,我们给出了线性、多项式、瓶颈型三种去噪器架构下生成样本分布的显式表达式。实证上,我们展示对主流UNet架构进行微小改动会导致生成结果的创造性形态发生显著变化,且常产生高度不真实的样本。综合来看,扩散模型的成功依赖于去噪器架构的归纳偏置与真实目标分布之间强匹配。

原文摘要 · Abstract (English)

The creativity of diffusion models refers to their ability to generate highly realistic images that are different from their training data. Creativity is somewhat surprising since it is known that if the denoiser used in the diffusion model is the Bayes optimal denoiser for a given training set, then the model will simply copy the training samples. In this paper we present empirical and theoretical results that suggest that creativity in diffusion models is due to an interaction between the denoiser architecture and the target distribution. Theoretically, we give explicit forms for the distribution of generated samples as a function of the target distribution and the denoiser architecture for three different denoiser architectures (linear, polynomial, bottleneck). Empirically, we show that small changes in the popular UNET denoiser architecture leads to very different forms of creativity, and these small changes often yield samples that are highly nonrealistic. Taken together, our results show that diffusion models will only be successful if the inductive bias of the denoiser architecture is in strong alignment with the true target distribution.

扩散模型去噪器创造力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。