提出统一框架,让一步扩散模型更高效且无需预训练。
On the Design of One-step Diffusion via Shortcutting Flow Paths
- 构建通用设计框架,解耦方法与实现细节。
- 一步生成达ImageNet上FID 2.85,两步达2.53。
- 无需预训练或教学蒸馏,适合组件级创新者。
近期的少步扩散模型通过捷径化概率路径提升了效率与效果,尤其在从零训练一步扩散模型(即捷径模型)方面表现突出。然而,其理论推导与实际实现常紧密耦合,模糊了设计空间。为此,我们提出一个适用于代表性捷径模型的通用设计框架,为模型有效性提供理论支持,并解耦具体组件选择,从而实现系统性改进。基于该框架,所提模型在无分类器引导设置下,于ImageNet-256x256上实现一步生成新最优FID50k 2.85,训练步数增至两倍时进一步达到FID50k 2.53。显著的是,模型无需预训练、蒸馏或课程学习。我们认为本工作降低了捷径模型组件级创新的门槛,推动其设计空间的有原则探索。
原文摘要 · Abstract (English)
Recent advances in few-step diffusion models have demonstrated their efficiency and effectiveness by shortcutting the probabilistic paths of diffusion models, especially in training one-step diffusion models from scratch (\emph{a.k.a.} shortcut models). However, their theoretical derivation and practical implementation are often closely coupled, which obscures the design space. To address this, we propose a common design framework for representative shortcut models. This framework provides theoretical justification for their validity and disentangles concrete component-level choices, thereby enabling systematic identification of improvements. With our proposed improvements, the resulting one-step model achieves a new state-of-the-art FID50k of 2.85 on ImageNet-256x256 under the classifier-free guidance setting with one step generation, and further reaches FID50k of 2.53 with 2x training steps. Remarkably, the model requires no pre-training, distillation, or curriculum learning. We believe our work lowers the barrier to component-level innovation in shortcut models and facilitates principled exploration of their design space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。