arXiv:2412.05984cs.CV2024-12CVPR被引 2

分层扩散模型通过语义层级生成,显著提升复杂图像质量。

Nested Diffusion Models Using Hierarchical Latent Priors

  • 用多级扩散模型逐层生成语义特征,上层控制下层
  • 在多个数据集上,图像质量显著优于基线,尤其无条件生成更强
  • 计算开销小,适合需要高质量图像生成的场景

我们提出嵌套扩散模型,一种高效且强大的层次化生成框架,显著提升扩散模型在复杂场景图像生成中的质量。该方法通过一系列扩散模型逐层生成不同语义层级的潜在变量,每个模型基于前一层更高层级模型的输出进行条件生成,最终完成图像生成。层次化潜在变量沿着预设的语义路径引导生成过程,使模型能捕捉精细结构细节并大幅提高图像质量。为构建这些潜在变量,我们利用预训练视觉编码器学习强语义表征,并通过降维和加噪调节其容量。在多个数据集上,系统在无条件与类别/文本条件生成任务中均表现出显著的质量提升。此外,我们的无条件生成系统明显优于基线条件系统。由于高层级使用低维表示,整体计算开销极小。

原文摘要 · Abstract (English)

We introduce nested diffusion models, an efficient and powerful hierarchical generative framework that substantially enhances the generation quality of diffusion models, particularly for images of complex scenes. Our approach employs a series of diffusion models to progressively generate latent variables at different semantic levels. Each model in this series is conditioned on the output of the preceding higher-level models, culminating in image generation. Hierarchical latent variables guide the generation process along predefined semantic pathways, allowing our approach to capture intricate structural details while significantly improving image quality. To construct these latent variables, we leverage a pre-trained visual encoder, which learns strong semantic visual representations, and modulate its capacity via dimensionality reduction and noise injection. Across multiple datasets, our system demonstrates significant enhancements in image quality for both unconditional and class/text conditional generation. Moreover, our unconditional generation system substantially outperforms the baseline conditional system. These advancements incur minimal computational overhead as the more abstract levels of our hierarchy work with lower-dimensional representations.

扩散模型图像生成层次结构潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。