提出连续生成新框架,提升离散序列生成质量与灵活性。
Discrete Stochastic Localization for Non-autoregressive Generation
- 用单位球面嵌入实现与信噪比无关的去噪机制
- 同一模型支持从128到1024步的多种采样路径
- 兼容随机顺序自回归采样,仅需48步混合采样
连续扩散是非自回归生成的自然框架,但在离散序列生成上通常落后于掩码离散扩散模型(MDMs)。我们指出瓶颈并非连续性本身,而是去噪依赖于时间步索引的噪声分布。本文提出离散随机定位(DSL),一种基于单位球面词元嵌入的连续状态框架,其贝叶斯最优去噪器在定位信道下对名义信噪比(SNR)不变。一个训练好的网络可支持整个按词元定义的SNR路径族,其中端点为掩码扩散路径。在OpenWebText数据集上,使用预训练的MDLM检查点进行DSL微调后,在所有步数预算(T=128至T=1024)下均显著提升分布保真度(MAUVE)。同一模型还可支持随机顺序自回归采样,以及仅需T=48步的连续-离散混合采样,无需蒸馏或重新训练。
原文摘要 · Abstract (English)
Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。