用离散潜在码提升扩散模型生成质量与多样性
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- 提出离散潜在码(DLC),用离散令牌序列替代连续嵌入
- 在ImageNet上实现无条件生成新纪录,且支持分布外样本生成
- 可组合生成新图像,适配文本到图像生成任务
我们认为扩散模型在建模复杂分布上的成功,主要源于其输入条件。本文从理想表示应提升生成保真度、易生成且具备组合性以生成训练分布外样本的角度,引入离散潜在码(DLC)。DLC基于自监督学习的单纯形嵌入训练得到,是离散令牌序列,而非标准连续图像嵌入。其易于生成且组合性支持生成超越训练分布的新图像。使用DLC训练的扩散模型在ImageNet上实现了无条件图像生成的新状态。此外,通过组合DLC可生成语义上新颖且连贯的图像,展现出多样化组合能力。最后,我们通过微调预训练文本扩散语言模型,使其实现文本到图像生成,直接生成能产生分布外新样本的DLC。
原文摘要 · Abstract (English)
We argue that diffusion models' success in modeling complex distributions is, for the most part, coming from their input conditioning. This paper investigates the representation used to condition diffusion models from the perspective that ideal representations should improve sample fidelity, be easy to generate, and be compositional to allow out-of-training samples generation. We introduce Discrete Latent Code (DLC), an image representation derived from Simplicial Embeddings trained with a self-supervised learning objective. DLCs are sequences of discrete tokens, as opposed to the standard continuous image embeddings. They are easy to generate and their compositionality enables sampling of novel images beyond the training distribution. Diffusion models trained with DLCs have improved generation fidelity, establishing a new state-of-the-art for unconditional image generation on ImageNet. Additionally, we show that composing DLCs allows the image generator to produce out-of-distribution samples that coherently combine the semantics of images in diverse ways. Finally, we showcase how DLCs can enable text-to-image generation by leveraging large-scale pretrained language models. We efficiently finetune a text diffusion language model to generate DLCs that produce novel samples outside of the image generator training distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。