arXiv:2509.23937cs.LGcond-mat.stat-mech2025-09被引 1

揭示扩散模型中信息分离机制,解释为何生成图像时语义与细节可独立控制。

On the Separability of Information in Diffusion Models

  • 发现模型大部分信息用于重建图像细节,而非语义。
  • 图像类别相关性由语义内容决定,与低层细节无关。
  • 解释无分类器引导为何能先影响语义再填充细节。

扩散模型通过在训练过程中将信息编码于神经网络,将噪声逐步转化为数据。本文探究了这些信息的本质:在像素空间的扩散模型中,(1)网络中大量信息用于重建图像的细微感知细节;(2)图像与其类别标签之间的关联由图像的语义内容决定,对低层细节不敏感。我们认为这些特性与数据本身的流形结构密切相关。最后,我们证明这些发现能解释无分类器引导的有效性:引导向量在生成初期放大图像与条件信号间的互信息,影响语义结构,但在后期细节填充阶段逐渐减弱。

原文摘要 · Abstract (English)

Diffusion models transform noise into data by injecting information that was captured in their neural network during the training phase. In this paper, we ask: \textit{what} is this information? We find that, in pixel-space diffusion models, (1) a large fraction of the total information in the neural network is committed to reconstructing small-scale perceptual details of the image, and (2) the correlations between images and their class labels are informed by the semantic content of the images, and are largely agnostic to the low-level details. We argue that these properties are intrinsically tied to the manifold structure of the data itself. Finally, we show that these facts explain the efficacy of classifier-free guidance: the guidance vector amplifies the mutual information between images and conditioning signals early in the generative process, influencing semantic structure, but tapers out as perceptual details are filled in.

扩散模型信息分离生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。