arXiv:2410.11081cs.LGstat.ML2024-10被引 273

提出稳定训练大模型的连续时间一致性生成方法

Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

  • 统一扩散模型理论框架,发现训练不稳定的根源
  • 仅用两步采样即达先进生成质量,多数据集FID领先
  • 适合追求高效生成与大规模训练的研究者

一致性模型(CMs)是一类基于扩散的生成模型,以快速采样为优势。现有大部分CMs采用离散时间步训练,引入额外超参数且易受离散化误差影响。虽然连续时间形式可缓解这些问题,但其成功受限于训练不稳定。本文提出一个简化的理论框架,统一了以往扩散模型与CMs的参数化方式,揭示了不稳定的本质原因。基于此分析,我们在扩散过程参数化、网络架构和训练目标上引入关键改进,使连续时间CMs得以在前所未有的规模下训练,达到ImageNet 512x512上15亿参数。所提训练算法仅用两个采样步骤,即在CIFAR-10上取得2.06的FID,ImageNet 64x64上为1.48,ImageNet 512x512上为1.88,与最优扩散模型的FID差距缩小至10%以内。

原文摘要 · Abstract (English)

Consistency models (CMs) are a powerful class of diffusion-based generative models optimized for fast sampling. Most existing CMs are trained using discretized timesteps, which introduce additional hyperparameters and are prone to discretization errors. While continuous-time formulations can mitigate these issues, their success has been limited by training instability. To address this, we propose a simplified theoretical framework that unifies previous parameterizations of diffusion models and CMs, identifying the root causes of instability. Based on this analysis, we introduce key improvements in diffusion process parameterization, network architecture, and training objectives. These changes enable us to train continuous-time CMs at an unprecedented scale, reaching 1.5B parameters on ImageNet 512x512. Our proposed training algorithm, using only two sampling steps, achieves FID scores of 2.06 on CIFAR-10, 1.48 on ImageNet 64x64, and 1.88 on ImageNet 512x512, narrowing the gap in FID scores with the best existing diffusion models to within 10%.

生成模型扩散模型连续时间高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。