提出新框架,让离散数据生成更高效且灵活。
Generalized Discrete Diffusion from Snapshots
- 用快照隐变量替代完整噪声路径,简化逆过程建模。
- 在大规模离散生成任务中训练更快、生成质量更高。
- 适合追求高效生成的科研与工程人员使用。
我们提出广义离散扩散从快照(GDDS),一种支持大离散状态空间中任意噪声过程的统一框架。该框架涵盖现有所有离散扩散方法,同时显著提升噪声动态选择的灵活性。前向噪声过程基于均匀化技术,实现快速任意污染。逆过程则基于快照隐变量推导出简洁的证据下界(ELBO),使标准生成模型架构可高效训练,并具备清晰的概率解释。在大规模词汇量离散生成任务上的实验表明,该框架在训练效率和生成质量上优于现有离散扩散方法,并首次在该规模上超越自回归模型。代码与项目博客见:https://oussamazekri.fr/gdds。
原文摘要 · Abstract (English)
We introduce Generalized Discrete Diffusion from Snapshots (GDDS), a unified framework for discrete diffusion modeling that supports arbitrary noising processes over large discrete state spaces. Our formulation encompasses all existing discrete diffusion approaches, while allowing significantly greater flexibility in the choice of corruption dynamics. The forward noising process relies on uniformization and enables fast arbitrary corruption. For the reverse process, we derive a simple evidence lower bound (ELBO) based on snapshot latents, instead of the entire noising path, that allows efficient training of standard generative modeling architectures with clear probabilistic interpretation. Our experiments on large-vocabulary discrete generation tasks suggest that the proposed framework outperforms existing discrete diffusion methods in terms of training efficiency and generation quality, and beats autoregressive models for the first time at this scale. We provide the code along with a blog post on the project page : \href{https://oussamazekri.fr/gdds}{https://oussamazekri.fr/gdds}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。