将复杂生成模型压缩到手机级设备,实现快速高质量图像生成。
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- 通过分布匹配优化,高效蒸馏高精度生成流模型。
- 支持少步生成,移动端1秒内出图,内存占用降低60%。
- 适合移动开发、边缘计算及资源受限场景的AI应用。
我们提出SD3.5-Flash,一种高效的少步蒸馏框架,将高质量图像生成能力带入消费级设备。该方法通过重构的分布匹配目标,蒸馏计算成本高昂的修正流模型,特别适配少步生成。引入两项关键创新:时间步共享以降低梯度噪声,分时间步微调以增强提示对齐。结合文本编码器重构与专用量化等全流程优化,系统在不同硬件上实现快速生成与内存高效部署。大规模用户评估表明,其性能持续优于现有少步方法,真正实现先进生成式AI的普适化落地。
原文摘要 · Abstract (English)
We present SD3.5-Flash, an efficient few-step distillation framework that brings high-quality image generation to accessible consumer devices. Our approach distills computationally prohibitive rectified flow models through a reformulated distribution matching objective tailored specifically for few-step generation. We introduce two key innovations: "timestep sharing" to reduce gradient noise and "split-timestep fine-tuning" to improve prompt alignment. Combined with comprehensive pipeline optimizations like text encoder restructuring and specialized quantization, our system enables both rapid generation and memory-efficient deployment across different hardware configurations. This democratizes access across the full spectrum of devices, from mobile phones to desktop computers. Through extensive evaluation including large-scale user studies, we demonstrate that SD3.5-Flash consistently outperforms existing few-step methods, making advanced generative AI truly accessible for practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。