arXiv:2412.16906cs.CV2024-12中稿 · AAAI被引 2

让文本生成图像更快更稳,一步到位且质量不降。

Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation

  • 用自校正流蒸馏融合一致性模型与对抗训练
  • 在少步甚至一步采样下仍保持高质量图像生成
  • 适合追求高效生成的视觉生成研究者

流匹配已成为一种有前景的生成模型训练框架,相比基于扩散的方法,其训练更为简便,并展现出出色的实证性能。然而,该方法在采样过程中仍需大量函数评估。为解决这一问题,我们提出一种自校正流蒸馏方法,在流匹配框架中有效整合一致性模型与对抗训练。本工作首次实现少步与一步采样下生成质量的一致性。大量实验验证了方法的有效性,在CelebA-HQ数据集上及COCO数据集的零样本基准测试中均取得优越的定量与定性结果。代码已开源:https://github.com/hao-pt/SCFlow.git。

原文摘要 · Abstract (English)

Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To address these limitations, we introduce a self-corrected flow distillation method that effectively integrates consistency models and adversarial training within the flow-matching framework. This work is a pioneer in achieving consistent generation quality in both few-step and one-step sampling. Our extensive experiments validate the effectiveness of our method, yielding superior results both quantitatively and qualitatively on CelebA-HQ and zero-shot benchmarks on the COCO dataset. Our implementation is released at https://github.com/hao-pt/SCFlow.git.

文本生成图像流匹配高效生成一致性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。