统一了掩码、连续与混合扩散模型的理论框架,揭示其内在机制。
Sticky Jump Diffusions: A Unifying View of Masked, Continuous, and Hybrid Diffusion

- 提出连续时间马尔可夫过程SJD,通过流平衡推导反向跳跃规则。
- 仅用一个去噪分类器即可无仿真训练,准确估计得分与反向跳变率。
- 新设计的跨位置混合核提升性能,适用于图像、文本等多场景。
我们提出黏滞跳跃扩散(SJD),一种定义在ℝᵈ上的连续时间马尔可夫过程,其离散锚点为令牌嵌入。前向过程中,锚点以风险率释放质量,释放的质量在连续空间中扩散;时间反演将一个由得分驱动的SDE与一个黏滞跳跃核耦合,其跳跃率和目的地由前向过程的通量平衡决定。我们通过去噪风险匹配(Denoising Hazard Matching)从单一去噪分类器中估计得分与各锚点的反向风险,采用无需仿真的交叉熵训练。SJD可退化为掩码扩散、连续扩散和混合扩散三种情形。其反向过程解释了各类方法原本视为给定的特性:掩码扩散中掩码不携带源令牌证据,因每个锚点的解黏核坍缩至同一吸收点;连续扩散需终端投影,因其前向边际无原子,否则通量平衡无法产生反向跳跃;混合扩散的更新规则(承诺率、目的地、漂移)均由通量平衡导出,而非独立设计。超越这些极限后,解黏核成为可设计空间:跨位置混合使每位置向邻域干净值或嵌入的组合腐化,将空间局部性或约束图等依赖结构转化为腐蚀本身的归纳偏置,在CIFAR-10、Text8和Sudoku上优于恒等核混合模型。
原文摘要 · Abstract (English)
We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete anchors are token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal couples a score-driven SDE with a sticky jump kernel whose rate and destination are fixed by flux balance with the forward law. We estimate the score and the per-anchor reverse hazards from a single denoising classifier via Denoising Hazard Matching, the hazard analogue of denoising score matching, with simulation-free cross-entropy training. SJD recovers masked diffusion, continuous diffusion, and hybrid diffusion as limits. Its reversal explains features that each family treats as given: the mask of masked diffusion carries no evidence about the source token because the unsticking kernel of every anchor collapses to the same absorbing point; the terminal projection of continuous diffusion is required due to the absence of atoms in its forward marginal, without which flux balance yields no reverse jumps; and the update rules of hybrid diffusion (commit rate, destination, and drift) all follow from flux balance rather than from separate design. Beyond these limits, the unsticking kernel becomes a design space: a cross-position blending corrupts each position toward a blend of its neighbors' clean values or embeddings, turning dependency structure such as spatial locality or a constraint graph into an inductive bias of the corruption itself, and improves over the identity-kernel hybrid on CIFAR-10, Text8, and Sudoku.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。