arXiv:2505.17384cs.LGcs.CV2025-05被引 12

提升离散扩散模型生成质量,尤其在少步数下表现更优

Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling

  • 引入隐变量建模增强维度间相关性捕捉
  • 少步数下生成样本质量显著优于基线
  • 适合需要高效高质量生成的文本与图像任务

离散扩散模型在建模复杂离散数据方面展现出巨大潜力,掩码扩散模型(MDMs)在生成质量与速度间取得良好平衡。MDMs 通过逐步解掩码多个维度从全掩码输入中去噪,但在较少去噪步数下,因维度间依赖关系建模不足导致性能下降。本文提出变分自编码离散扩散(VADD),通过引入辅助识别模型,在潜变量空间中隐式捕捉维度间相关性。该方法利用变分下界最大化实现稳定训练,并支持对训练集的推理。在保留传统 MDM 高效性的基础上,显著提升样本质量,尤其在去噪步数较少时表现突出。在二维玩具数据、像素级图像生成和文本生成任务上的实验表明,VADD 在少步数下持续优于 MDM 基线。

原文摘要 · Abstract (English)

Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively unmasking multiple dimensions from an all-masked input, but their performance can degrade when using few denoising steps due to limited modeling of inter-dimensional dependencies. In this paper, we propose Variational Autoencoding Discrete Diffusion (VADD), a novel framework that enhances discrete diffusion with latent variable modeling to implicitly capture correlations among dimensions. By introducing an auxiliary recognition model, VADD enables stable training via variational lower bounds maximization and amortized inference over the training set. Our approach retains the efficiency of traditional MDMs while significantly improving sample quality, especially when the number of denoising steps is small. Empirical results on 2D toy data, pixel-level image generation, and text generation demonstrate that VADD consistently outperforms MDM baselines in sample quality with few denoising steps.

离散扩散生成模型隐变量建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。