提出连续生成新框架,提升离散序列生成质量与灵活性。
Discrete Stochastic Localization for Non-autoregressive Generation
- 用单位球面嵌入实现与信噪比无关的去噪机制
- 单个模型支持多种采样路径,128到1024步均更保真
- 兼容随机顺序自回归与混合采样,无需重训
连续扩散是无自回归生成的自然框架,但在离散序列生成上通常落后于掩码离散扩散模型(MDMs)。我们指出瓶颈并非连续性本身,而是去噪依赖于时间步索引的噪声模式。本文提出离散随机定位(DSL),一种基于单位球面词元嵌入的连续状态框架,其贝叶斯最优去噪器在定位信道下对名义信噪比(SNR)保持不变。一个训练好的网络即可支持整个每词元SNR路径族,掩码扩散路径为其特例。在OpenWebText上,用预训练的MDLM检查点进行微调后,所有步数预算从T=128到T=1024均显著提升分布保真度(MAUVE);同一检查点还支持随机顺序自回归采样及仅需T=48步的连续-离散混合采样,无需蒸馏或重新训练。
原文摘要 · Abstract (English)
Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。