让大模型几步生成,几乎零成本。
Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- 用自蒸馏让流匹配模型跳步采样,无需重训
- 3步生成仅需不到一天的算力(1 A100)
- 支持少样本微调,适合高效部署场景
我们提出一种超高效的后训练方法,可将大规模预训练流匹配扩散模型压缩为高效少步采样器,核心是新型速度场自蒸馏机制。传统跳步采样虽灵活,但依赖专用步长嵌入,与现有模型不兼容,除非从头重训——代价近似于预训练本身。本文通过在速度场而非样本空间上进行在线自引导蒸馏,使标准流匹配模型(如Flux)可直接获得更激进的跳步能力,无需步长嵌入。该方法训练高效,例如3步版Flux可在不到一天内完成(1 A100)。此外,该方法可融入预训练阶段,使模型天生具备高质量少步生成能力。同时,首次实现对十亿参数级扩散模型的少样本蒸馏(如仅10组图文对),性能达顶尖水平,成本几乎可忽略。
原文摘要 · Abstract (English)
We present an ultra-efficient post-training method for shortcutting large-scale pre-trained flow matching diffusion models into efficient few-step samplers, enabled by novel velocity field self-distillation. While shortcutting in flow matching, originally introduced by shortcut models, offers flexible trajectory-skipping capabilities, it requires a specialized step-size embedding incompatible with existing models unless retraining from scratch$\unicode{x2013}$a process nearly as costly as pretraining itself. Our key contribution is thus imparting a more aggressive shortcut mechanism to standard flow matching models (e.g., Flux), leveraging a unique distillation principle that obviates the need for step-size embedding. Working on the velocity field rather than sample space and learning rapidly from self-guided distillation in an online manner, our approach trains efficiently, e.g., producing a 3-step Flux less than one A100 day. Beyond distillation, our method can be incorporated into the pretraining stage itself, yielding models that inherently learn efficient, few-step flows without compromising quality. This capability also enables, to our knowledge, the first few-shot distillation method (e.g., 10 text-image pairs) for dozen-billion-parameter diffusion models, delivering state-of-the-art performance at almost free cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。