arXiv:2510.25544stat.MLcs.IT2025-10被引 9

提出掩码扩散模型的误差界与最优生成调度,兼顾效率与精度。

Error Bounds and Optimal Schedules for Masked Diffusions with Factorized Approximations

  • 基于条件独立近似降低计算开销,用平均每步生成词数控制误差。
  • 证明误差仅与平均生成量相关,与序列长度无关,解释模型成功原因。
  • 发现最优调度依赖数据信息分布,可指导高效采样策略设计。

近期提出的离散数据生成模型,如掩码扩散模型(MDMs),利用条件独立近似来降低经典自回归模型(ARMs)的计算成本,代价是采样分布存在一定偏差。本文研究由此产生的计算与精度权衡,给出了仅依赖每迭代平均生成词数的相对熵误差上界,且与数据维度(即序列长度)无关,从而支持了MDMs的实证成功。进一步分析非恒定调度(即生成过程中动态调整未掩码词数)带来的收益,识别出最优调度作为数据分布信息轮廓的函数,实现调度大小的合理优化。方法直接定义为采样算法,不依赖经典的时间反演扩散过程推导,使证明更简洁透明。

原文摘要 · Abstract (English)

Recently proposed generative models for discrete data, such as Masked Diffusion Models (MDMs), exploit conditional independence approximations to reduce the computational cost of popular Auto-Regressive Models (ARMs), at the price of some bias in the sampling distribution. We study the resulting computation-vs-accuracy trade-off, providing general error bounds (in relative entropy) that depend only on the average number of tokens generated per iteration and are independent of the data dimensionality (i.e. sequence length), thus supporting the empirical success of MDMs. We then investigate the gain obtained by using non-constant schedule sizes (i.e. varying the number of unmasked tokens during the generation process) and identify the optimal schedule as a function of a so-called information profile of the data distribution, thus allowing for a principled optimization of schedule sizes. We define methods directly as sampling algorithms and do not use classical derivations as time-reversed diffusion processes, leading us to simple and transparent proofs.

生成模型扩散模型优化调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。