提出新条件机制,显著提升扩散模型生成质量与效率
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
- 分离语义与控制信息的条件机制,提升模型灵活性
- 在ImageNet-1k和CC12M上分别实现7%-23%的FID改进
- 公开复现5个模型,推动可复现性研究
大规模潜在扩散模型(LDM)训练已实现图像生成质量的飞跃。然而,表现最佳的LDM训练方案关键组件常未公开,阻碍了领域进展的验证与公平比较。本文深入研究LDM训练方法,聚焦模型性能与训练效率。为确保公平对比,我们复现了五篇已有论文的模型及其训练配方。研究发现:(i) 条件机制对文本提示等语义信息及裁剪尺寸、随机翻转等控制元数据的处理方式显著影响模型性能;(ii) 在小规模低分辨率数据集上学习的表征迁移到大规模数据集,可提升训练效率与模型表现。基于此,我们提出一种新条件机制,解耦语义与控制信息条件,实现在ImageNet-1k上的类别条件生成新基准:256×256分辨率下FID降低7%,512×512分辨率下降低8%;在CC12M数据集上,文本到图像生成的FID在256×256分辨率下降低8%,512×512分辨率下降低23%。
原文摘要 · Abstract (English)
Large-scale training of latent diffusion models (LDMs) has enabled unprecedented quality in image generation. However, the key components of the best performing LDM training recipes are oftentimes not available to the research community, preventing apple-to-apple comparisons and hindering the validation of progress in the field. In this work, we perform an in-depth study of LDM training recipes focusing on the performance of models and their training efficiency. To ensure apple-to-apple comparisons, we re-implement five previously published models with their corresponding recipes. Through our study, we explore the effects of (i)~the mechanisms used to condition the generative model on semantic information (e.g., text prompt) and control metadata (e.g., crop size, random flip flag, etc.) on the model performance, and (ii)~the transfer of the representations learned on smaller and lower-resolution datasets to larger ones on the training efficiency and model performance. We then propose a novel conditioning mechanism that disentangles semantic and control metadata conditionings and sets a new state-of-the-art in class-conditional generation on the ImageNet-1k dataset -- with FID improvements of 7% on 256 and 8% on 512 resolutions -- as well as text-to-image generation on the CC12M dataset -- with FID improvements of 8% on 256 and 23% on 512 resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。