提出可自适应采样的离散扩散模型,突破维度依赖瓶颈。
Provably adaptive sampling with uniform and remasking discrete diffusion models
- 设计基于留一法去噪器的并行采样器,支持均匀与重掩码过程
- 采样复杂度仅依赖目标分布内在相关性(DTC),不随维度线性增长
- 理论证明采样误差由得分估计与离散化误差共同决定,适合高维结构数据
离散扩散模型通过并行更新提供对自回归生成的有前途替代方案,但其采样效率高度依赖前向过程和采样器选择。针对均匀前向过程,现有标准τ-跳跃采样器的下界随环境维度d线性增长,引发该依赖是否为前向过程固有属性的疑问。本文回答否定:考虑一种基于留一法去噪器的一阶采样器,适用于均匀与重掩码过程,其坐标更新可并行进行。在两种情况下,采样过程可纠正去噪错误,当大量坐标同时更新时尤为必要。主要结果建立自适应采样保证:在对数因子范围内,N = O(DTC(X₀)/ε)个离散步骤足以实现采样误差O(ε_score + ε),其中ε_score为得分估计误差。因此,采样复杂度由目标分布的双重总相关性DTC(X₀)决定,而非直接依赖于环境维度d。分析通过贝叶斯最优辅助采样器完成,将离散化误差与得分估计误差分离。还推导出离散化误差的精确信息论表达式,涉及前向过程中不同时间点各坐标间的互信息。该表达式适用于一般前向过程,在均匀与重掩码情况下可由DTC(X₀)控制。数值实验在结构化合成分布上验证了预测的维度自适应行为。
原文摘要 · Abstract (English)
Discrete diffusion models offer a promising alternative to autoregressive generation by enabling parallel updates, but their sampling efficiency can depend strongly on the choice of the forward process and the sampler. For the uniform forward process, existing lower bounds for the standard $τ$-leaping sampler scale linearly with the ambient dimension $d$, raising the question of whether this dependence is intrinsic to the forward process. We answer this question in the negative. We consider a first-order sampler based on the leave-one-out denoiser for uniform and remasking processes whose coordinate updates can be performed in parallel. In both cases, the sampler can correct denoising mistakes during the sampling process, which becomes necessary when many coordinates are updated together. Our main result establishes an adaptive sampling guarantee: up to logarithmic factors, $N = O(\mathrm{DTC}(X_0) / \varepsilon)$ discretization steps suffice to achieve sampling error $O(\varepsilon_{\mathrm{score}}+\varepsilon)$, where $\varepsilon_{\mathrm{score}}$ is the error in score estimation. Thus, the sampling complexity is governed by the intrinsic dependence structure of the target distribution, as measured by its dual total correlation $\mathrm{DTC}(X_0)$, rather than directly by the ambient dimension $d$. Our analysis proceeds through a Bayes-optimal auxiliary sampler that separates discretization error from score-estimation error. We also derive an exact information-theoretic representation of the discretization error in terms of the mutual information between different coordinates of the forward process at different times. This representation applies to general forward processes and, in the uniform and remasking cases, can be controlled by $\mathrm{DTC}(X_0)$. Numerical experiments on structured synthetic distributions illustrate the predicted dimension-adaptive behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。