首次为吸收型离散扩散模型提供收敛性保证,解决生成质量敏感问题。
Absorb and Converge: Provable Convergence Guarantee for Absorbing Discrete Diffusion Models
- 引入代理初始化分布,突破吸收态分布导致KL发散的理论瓶颈。
- 证明吸收率矩阵下采样器收敛速度优于均匀率矩阵,且无需提前停止。
- 提出新工具包:凸分析界、吸收得分函数控制、初始阶段非发散上界。
离散状态空间扩散模型在文本与图像生成等离散数据任务中表现优异,其性能对速率矩阵选择极为敏感,尤其在均匀与吸收型速率矩阵之间差异显著。尽管实证表明吸收型矩阵常带来更优生成质量,但现有理论研究主要集中在均匀矩阵情形,缺乏对吸收型扩散模型的收敛性保证与误差分析。本文首次为使用吸收速率矩阵的离散扩散模型提供有限时间误差界与收敛速率分析。通过构造代理初始化分布,克服吸收态稳态分布为单点导致的KL散度未定义问题,推导前向过程的KL上界。进一步,针对τ-跃迁与统一化采样器,建立首例基于吸收速率矩阵的收敛性保障,并在合理假设下实现无需早停的收敛结果。分析中引入多项新技术:用于前向过程的Jensen型论证、吸收得分函数的新界方法,以及初始阶段得分函数的非发散上界,彻底消除早停依赖。
原文摘要 · Abstract (English)
Discrete state space diffusion models have shown significant advantages in applications involving discrete data, such as text and image generation. It has also been observed that their performance is highly sensitive to the choice of rate matrices, particularly between uniform and absorbing rate matrices. While empirical results suggest that absorbing rate matrices often yield better generation quality compared to uniform rate matrices, existing theoretical works have largely focused on the uniform rate matrices case. Notably, convergence guarantees and error analyses for absorbing diffusion models are still missing. In this work, we provide the first finite-time error bounds and convergence rate analysis for discrete diffusion models using absorbing rate matrices. We begin by deriving an upper bound on the KL divergence of the forward process, introducing a surrogate initialization distribution to address the challenge posed by the absorbing stationary distribution, which is a singleton and causes the KL divergence to be ill-defined. We then establish the first convergence guarantees for both the $τ$-leaping and uniformization samplers under absorbing rate matrices, demonstrating improved rates over their counterparts using uniform rate matrices. Furthermore, under suitable assumptions, we provide convergence guarantees without early stopping. Our analysis introduces several new technical tools to address challenges unique to absorbing rate matrices. These include a Jensen-type argument for bounding forward process convergence, novel techniques for bounding absorbing score functions, and a non-divergent upper bound on the score near initialization that removes the need of early-stopping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。