arXiv:2508.14413cs.LGcs.CV2025-08

用极少隐状态训练扩散模型,速度提升4-6倍且效果不降。

Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states

  • 通过优化噪声调度,在32个隐状态上实现与1000个状态相当的效果。
  • 将隐状态数压缩至1个,仍可生成高质量图像。
  • 多模型组合实现快速分布式训练,适合大规模生成任务。

我们挑战了扩散模型的一个基本假设:训练需大量隐状态或时间步以使反向生成过程接近高斯分布。首先,我们证明通过精心设计的噪声调度,仅需约32个隐状态即可达到与约1000个隐状态训练模型相当的性能。其次,我们将该极限推进至单个隐状态,即T空间中的完全解耦。我们发现,通过组合多个独立训练的单隐状态模型,可轻松生成高质量样本。大量实验证明,所提出的解耦模型在两个不同数据集上,各项指标的收敛速度均提升了4-6倍。

原文摘要 · Abstract (English)

We challenge a fundamental assumption of diffusion models, namely, that a large number of latent-states or time-steps is required for training so that the reverse generative process is close to a Gaussian. We first show that with careful selection of a noise schedule, diffusion models trained over a small number of latent states (i.e. $T \sim 32$) match the performance of models trained over a much large number of latent states ($T \sim 1,000$). Second, we push this limit (on the minimum number of latent states required) to a single latent-state, which we refer to as complete disentanglement in T-space. We show that high quality samples can be easily generated by the disentangled model obtained by combining several independently trained single latent-state models. We provide extensive experiments to show that the proposed disentangled model provides 4-6$\times$ faster convergence measured across a variety of metrics on two different datasets.

扩散模型加速训练隐状态压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。