arXiv:2410.02667cs.LGhep-th2024-10被引 4

提出统一扩散框架,融合多种生成模型优势。

GUD: Generation with Unified Diffusion

  • 在任意基底下统一扩散与自回归模型,支持灵活表示与噪声调度。
  • 通过软条件机制平滑衔接扩散与自回归,实现跨架构融合。
  • 为高效训练和新生成任务设计提供广阔探索空间,适合模型架构研究者。

扩散生成模型通过反转逐步加噪的过程将噪声转化为数据。受物理学中重整化群思想启发,我们重新审视扩散模型,聚焦三个关键设计维度:1)扩散过程所用的表示形式(如像素、PCA、傅里叶或小波基);2)数据在扩散过程中转换的先验分布(如协方差为Σ的高斯分布);3)对数据不同部分分别施加噪声水平的分量级噪声调度。结合这些设计灵活性,我们构建了一个统一的扩散生成模型框架,显著提升设计自由度。特别地,我们引入软条件模型,可在任意基底下平滑插值标准扩散模型与自回归模型,概念上弥合两者差异。该框架拓展了广泛的设计空间,有望实现更高效的训练与生成,并为融合不同生成方法及任务的新架构铺平道路。

原文摘要 · Abstract (English)

Diffusion generative models transform noise into data by inverting a process that progressively adds noise to data samples. Inspired by concepts from the renormalization group in physics, which analyzes systems across different scales, we revisit diffusion models by exploring three key design aspects: 1) the choice of representation in which the diffusion process operates (e.g. pixel-, PCA-, Fourier-, or wavelet-basis), 2) the prior distribution that data is transformed into during diffusion (e.g. Gaussian with covariance $Σ$), and 3) the scheduling of noise levels applied separately to different parts of the data, captured by a component-wise noise schedule. Incorporating the flexibility in these choices, we develop a unified framework for diffusion generative models with greatly enhanced design freedom. In particular, we introduce soft-conditioning models that smoothly interpolate between standard diffusion models and autoregressive models (in any basis), conceptually bridging these two approaches. Our framework opens up a wide design space which may lead to more efficient training and data generation, and paves the way to novel architectures integrating different generative approaches and generation tasks.

扩散模型生成模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。