离散扩散模型先学数据支撑结构,再学频率细节。
Support Before Frequency in Discrete Diffusion

- 通过反向编辑分解出支持层与频率层的独立学习机制。
- 支持结构恢复早于频率排序,在小噪声阶段显著提前。
- 适用于语言建模、生成任务中的模型理解与设计。
离散扩散模型在语言建模中表现日益优异,但其去噪目标如何组织学习仍不清晰。尽管这些目标针对完整数据分布,我们证明在小噪声条件下,精确反向过程会形成粗粒度支持信息与细粒度频率信息的层级结构。对于均匀和吸收型(即掩码型)扩散,单个词元的反向编辑可分解为决定是否向数据支持(如语法正确句)靠近的主导尺度,以及同一尺度内相对概率的精细系数。因此,恢复有效性结构只需学习逆概率的数量级,而恢复数据频率需精细系数估计。该分离机制依赖于具体方法:均匀扩散呈现三类编辑(改善、保持、恶化有效性),而吸收扩散将主导质量集中于有效性改善动作。在掩码语言扩散模型与合成正则语言任务上的实验支持上述预测:支持定位早于支持内频率排序,且均匀与吸收扩散的速率差异符合预期。结果表明,离散扩散模型先学习数据支持,后学习数据频率。
原文摘要 · Abstract (English)
Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives target the full data distribution, we show that the exact reverse process induces a hierarchy between coarse support information and finer frequency information. For uniform and absorbing (a.k.a. masking) diffusion, we prove that, in the small-noise regime of the final denoising steps, each single-token reverse edit decomposes into a leading scale, determined by whether it moves toward the data support (e.g., grammatically valid sentences), and a finer coefficient, determining relative probabilities within the same scale. Thus, recovering validity structure only requires learning the correct order of magnitude of reverse probabilities, whereas recovering data frequencies requires coefficient-level estimation. The separation is mechanism-dependent: uniform diffusion exhibits a trichotomy into validity-improving, validity-preserving, and validity-worsening edits, while absorbing diffusion places its leading-order mass on validity-improving moves. Experiments on a masked language diffusion model and synthetic regular-language tasks support these predictions: support-localization emerges earlier than within-support frequency ranking, and the contrast between uniform and absorbing diffusion matches the predicted rate separation. Together, our results suggest that discrete diffusion models learn data support before data frequencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。