arXiv:2510.01329stat.MLcs.LG2025-10被引 26

用连续隐空间增强离散扩散模型,让掩码信息不丢失,生成更高质量结果。

Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling

  • 引入连续隐空间作为掩码的语义提示,避免信息真空。
  • 在文本、图像、代码生成中均提升质量,优于现有离散扩散模型。
  • 可灵活调节生成多样性与准确性,适合不同应用场景。

标准离散扩散模型将所有未观测状态统一映射到一个吸收态[MASK],导致去噪过程中语义信息丢失,形成‘信息空洞’。本文提出连续增强的离散扩散模型(CADD),在离散状态空间外引入一个连续潜在空间的配对扩散过程。该设计使掩码项以带有噪声但仍具信息量的连续向量表示,而非完全失真的[MASK]。在反向去噪步骤中,可利用连续潜在向量提供语义提示,引导离散重构。该框架结构简洁,兼容现有离散扩散训练方式。采样时,通过调整连续潜在向量估计器的强度与选择,可在模式覆盖(生成多样输出)与模式聚焦(生成上下文精准输出)之间实现可控权衡。实验表明,CADD在文本生成、图像合成和代码建模任务中均显著提升生成质量,在定性与定量指标上持续优于强基准模型。

原文摘要 · Abstract (English)

Standard discrete diffusion models treat all unobserved states identically by mapping them to an absorbing [MASK] token. This creates an 'information void' where semantic information that could be inferred from unmasked tokens is lost between denoising steps. We introduce Continuously Augmented Discrete Diffusion (CADD), a framework that augments the discrete state space with a paired diffusion in a continuous latent space. This yields graded, gradually corrupted states in which masked tokens are represented by noisy yet informative latent vectors rather than collapsed 'information voids'. At each reverse step, CADD may leverage the continuous latent as a semantic hint to guide discrete denoising. The design is clean and compatible with existing discrete diffusion training. At sampling time, the strength and choice of estimator for the continuous latent vector enables a controlled trade-off between mode-coverage (generating diverse outputs) and mode-seeking (generating contextually precise outputs) behaviors. Empirically, we demonstrate CADD improves generative quality over mask-based diffusion across text generation, image synthesis, and code modeling, with consistent gains on both qualitative and quantitative metrics against strong discrete baselines.

扩散模型离散生成语义提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。