arXiv:2502.05743cs.LGcs.CV2025-02NeurIPS被引 18

发现扩散模型在中等噪声下特征表现最佳,与泛化能力直接相关。

Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling

  • 利用图像数据低维结构,理论解释特征动态呈单峰现象
  • 中等噪声时特征质量最高,噪声过高或过低则下降
  • 单峰动态可作为模型是否泛化的可靠指标

扩散模型虽以生成任务设计,却展现出出色的自监督表征学习能力。一个引人注目的现象是:特征质量在中等噪声水平时达到峰值,呈现单峰动态。本文从理论与实证两方面深入研究该现象。基于图像数据的低维特性,理论上证明当扩散模型成功捕捉数据分布时,单峰动态便会出现;其本质是去噪强度与类别置信度在不同噪声尺度下的协同作用。实验表明,在分类任务中,单峰动态的出现可靠反映模型的泛化能力:当模型能生成新图像时出现单峰;一旦开始记忆训练数据,动态即转为单调下降。

原文摘要 · Abstract (English)

Diffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal representation dynamics, where the quality of learned features peaks at an intermediate noise level. In this work, we conduct a comprehensive theoretical and empirical investigation of this phenomenon. Leveraging the inherent low-dimensionality structure of image data, we theoretically demonstrate that the unimodal dynamic emerges when the diffusion model successfully captures the underlying data distribution. The unimodality arises from an interplay between denoising strength and class confidence across noise scales. Empirically, we further show that, in classification tasks, the presence of unimodal dynamics reliably reflects the generalization of the diffusion model: it emerges when the model generates novel images and gradually transitions to a monotonically decreasing curve as the model begins to memorize the training data.

扩散模型表征学习单峰动态泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。