arXiv:2410.14758cs.LG2024-10被引 5

用一致性匹配提升离散向量量化图像生成的稳定性与质量

Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion

  • 将嵌入空间与连续潜变量扩散模型结合,实现端到端联合训练
  • 在ImageNet上仅用50步即达FID 6.81,优于现有离散潜变量方法
  • 引入一致性匹配损失与噪声调度策略,防止嵌入崩溃

通过将离散表示嵌入连续潜空间,可利用连续空间潜扩散模型处理离散数据的生成建模。然而,尽管初期取得成功,大多数潜扩散方法依赖固定预训练嵌入,限制了与扩散模型的联合训练优势。虽然联合学习嵌入(通过重构损失)与潜扩散模型(通过得分匹配损失)可提升性能,但端到端训练易引发嵌入崩溃,降低生成质量。为此,我们提出VQ-LCMD,一种嵌入空间内的连续潜扩散框架,可稳定训练。VQ-LCMD采用新训练目标,结合嵌入-扩散变分下界与一致性匹配(CM)损失,辅以偏移余弦噪声调度和随机丢弃策略。在多个基准测试中,所提VQ-LCMD在FFHQ、LSUN Churches和LSUN Bedrooms上表现更优。特别地,在ImageNet上,类条件图像生成仅用50步即达FID 6.81。

原文摘要 · Abstract (English)

By embedding discrete representations into a continuous latent space, we can leverage continuous-space latent diffusion models to handle generative modeling of discrete data. However, despite their initial success, most latent diffusion methods rely on fixed pretrained embeddings, limiting the benefits of joint training with the diffusion model. While jointly learning the embedding (via reconstruction loss) and the latent diffusion model (via score matching loss) could enhance performance, end-to-end training risks embedding collapse, degrading generation quality. To mitigate this issue, we introduce VQ-LCMD, a continuous-space latent diffusion framework within the embedding space that stabilizes training. VQ-LCMD uses a novel training objective combining the joint embedding-diffusion variational lower bound with a consistency-matching (CM) loss, alongside a shifted cosine noise schedule and random dropping strategy. Experiments on several benchmarks show that the proposed VQ-LCMD yields superior results on FFHQ, LSUN Churches, and LSUN Bedrooms compared to discrete-state latent diffusion models. In particular, VQ-LCMD achieves an FID of 6.81 for class-conditional image generation on ImageNet with 50 steps.

图像生成扩散模型向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。