用连续时间流映射实现高效少步图像生成,性能超越现有方法。
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
- 提出新型连续时间流映射训练目标,兼容多步采样。
- 在ImageNet 64x64和512x512上实现少步生成新纪录。
- 适合追求高效生成的模型部署与文本到图像应用。
扩散模型和流模型是当前最先进的生成建模方法,但需要大量采样步骤。一致性模型可将其压缩为一步生成器;然而,其性能随步数增加而下降,本文从理论和实证两方面证明了这一点。流映射通过单步连接任意噪声水平,在所有步数下均保持有效。本文提出两种新的连续时间目标用于训练流映射,并引入创新训练技巧,推广了现有的一致性与流匹配目标。进一步表明,使用低质量模型进行自引导可提升性能,对抗微调还能额外提升效果,且样本多样性损失极小。我们在多个挑战性图像生成基准上广泛验证了名为Align Your Flow的流映射模型,在ImageNet 64x64和512x512上均以小型高效神经网络实现了少步生成最优表现。最后,文本到图像流映射模型在非对抗训练的少步采样中优于所有现有方法。
原文摘要 · Abstract (English)
Diffusion- and flow-based models have emerged as state-of-the-art generative modeling approaches, but they require many sampling steps. Consistency models can distill these models into efficient one-step generators; however, unlike flow- and diffusion-based methods, their performance inevitably degrades when increasing the number of steps, which we show both analytically and empirically. Flow maps generalize these approaches by connecting any two noise levels in a single step and remain effective across all step counts. In this paper, we introduce two new continuous-time objectives for training flow maps, along with additional novel training techniques, generalizing existing consistency and flow matching objectives. We further demonstrate that autoguidance can improve performance, using a low-quality model for guidance during distillation, and an additional boost can be achieved by adversarial finetuning, with minimal loss in sample diversity. We extensively validate our flow map models, called Align Your Flow, on challenging image generation benchmarks and achieve state-of-the-art few-step generation performance on both ImageNet 64x64 and 512x512, using small and efficient neural networks. Finally, we show text-to-image flow map models that outperform all existing non-adversarially trained few-step samplers in text-conditioned synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。