arXiv:2505.19656cs.CV2025-05被引 3

改进离散扩散模型的噪声设计,提升生成质量与多样性。

ReDDiT: Rehashing Noise for Discrete Visual Generation

  • 引入随机多索引污染机制,扩展潜在变量的演化路径。
  • 生成质量显著提升,gFID从6.18降至1.61,接近连续模型水平。
  • 适合关注离散生成效率与稳定性的研究人员使用。

在视觉生成领域,离散扩散模型因其高效性和兼容性日益受到关注。然而,现有方法仍落后于连续模型,我们归因于噪声(吸收态)设计和采样启发式策略。本文提出一种针对离散扩散变压器的重哈希噪声方法(ReDDiT),旨在扩展吸收态并增强模型表达能力。ReDDiT通过随机多索引污染丰富训练期间潜在变量的演化路径。所提出的重哈希采样器可逆向随机吸收路径,确保生成过程高多样性且低偏差。这些重构使生成结果更一致、更具竞争力,减少了对复杂随机性调优的需求。实验表明,ReDDiT显著优于基线模型(gFID从6.18降至1.61),性能与连续模型相当。

原文摘要 · Abstract (English)

In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous counterparts, which we attribute to noise (absorbing state) design and sampling heuristics. In this study, we propose a rehashing noise approach for discrete diffusion transformer (termed ReDDiT), with the aim to extend absorbing states and improve expressive capacity of discrete diffusion models. ReDDiT enriches the potential paths that latent variables traverse during training with randomized multi-index corruption. The derived rehash sampler, which reverses the randomized absorbing paths, guarantees high diversity and low discrepancy of the generation process. These reformulations lead to more consistent and competitive generation quality, mitigating the need for heavily tuned randomness. Experiments show that ReDDiT significantly outperforms the baseline model (reducing gFID from 6.18 to 1.61) and is on par with the continuous counterparts.

离散生成扩散模型噪声设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。