提出新理论框架,让离散扩散模型在任意语言任务中收敛更快更稳定。
Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space
- 用伴随方程分析可观测变量,避开概率分布直接计算
- 在任意积分概率度量下实现无维度依赖的收敛保证
- 适用于掩码和均匀先验,适合大词汇量生成任务
离散扩散已成为语言、视觉和生物等领域的主流生成建模框架。但现有收敛理论存在根本局限:基于KL的分析在奇异先验(如掩码分布)下发散,总变差(TV)界依赖状态空间大小S,对现代语言任务(词汇量达数十万)已无意义。本文提出统一的伴随方程框架,在任意积分概率度量(IPM)下建立无维度依赖的收敛保证。据我们所知,这是首个完全摆脱S依赖且同时适用于掩码与均匀先验的理论结果。更重要的是,该理论可将已有步数复杂度保证推广至任意IPM。理论仅需单一标准速率矩阵正则性假设,适用于一般先验。五大创新技术推动突破:1. 通过伴随方程在可观测空间工作;2. 正则性分析导出任意IPM界;3. 耦合论证消除均匀转移下的S依赖;4. 评分边缘抵消;5. 退出路由技术消除掩码转移下的S依赖。本框架显著区别于以往路径空间KL与TV方法,避免其缺陷。除收敛界外,还为离散扩散模型提供理论研究工具,包括损失函数的合理选择与无维度步数复杂度分析。
原文摘要 · Abstract (English)
Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exhibits fundamental limitations. KL-based analyses diverge under singular priors such as the masked distribution, while bounds in total variation (TV) depend on the state space size $S$ and become vacuous for modern language tasks, where vocabularies contain hundreds of thousands of tokens. We develop a unified adjoint-equation-based framework that establishes dimension-free convergence guarantees in any integral probability metric (IPM). To the best of our knowledge, our bounds are the first to be entirely free of $S$ and applicable to both masked and uniform priors. Importantly, our results can extend existing step complexity guarantees to any IPM. In addition, our theory relies only on a single standard rate-matrix regularity assumption and applies to general priors. Five novel techniques drive our improvements: 1. working in the space of observables via adjoint equations rather than directly with probability measures; 2. a regularity analysis that yields bounds on any IPM; 3. a coupling argument that removes $S$-dependence under uniform transitions; and 4. score-marginal cancellation and 5. exit-routing techniques that remove $S$-dependence under masked transitions. Our framework thus sharply departs from prior analyses and avoids the shortcomings of pathspace-KL and existing TV-based approaches. Beyond convergence bounds, our framework provides a versatile toolkit for further theoretical study of discrete diffusion models, including principled choices of loss functions and dimension-free step complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。