解决生成模型中五项核心缺陷,提升单网络多步采样性能。
Improved Training Technique for Shortcut Models
- 提出统一训练框架iSM,动态控制引导强度,消除累积偏差。
- 在ImageNet上实现显著FID降低,支持一步、几步及多步采样。
- 适合需要高效生成且对细节要求高的研究者和开发者。
快捷模型代表了一种有前景的非对抗性生成建模范式,能通过单一训练网络实现一步、几步乃至多步采样。然而其广泛应用受限于五个关键性能瓶颈:(1)首次形式化了引导累积的隐藏缺陷,导致严重图像伪影;(2)固定引导方式缺乏推理时控制灵活性;(3)依赖直接域低层距离引发频率偏差,使重建偏向低频;(4)与EMA训练冲突导致自洽性发散;(5)轨迹弯曲阻碍收敛。为此,本文提出iSM统一训练框架,通过四项改进系统性解决上述问题:内在引导提供显式动态控制,缓解引导累积与僵化问题;多层级小波损失减轻频率偏差,恢复高频细节;缩放最优传输(sOT)降低训练方差,学习更直线稳定的生成路径;双EMA策略调和训练稳定与自洽性。ImageNet 256×256上的大量实验表明,该方法在一步、几步及多步生成中均显著优于基线快捷模型,使快捷模型成为可实用且具竞争力的生成模型类别。
原文摘要 · Abstract (English)
Shortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained network. However, their widespread adoption has been stymied by critical performance bottlenecks. This paper tackles the five core issues that held shortcut models back: (1) the hidden flaw of compounding guidance, which we are the first to formalize, causing severe image artifacts; (2) inflexible fixed guidance that restricts inference-time control; (3) a pervasive frequency bias driven by a reliance on low-level distances in the direct domain, which biases reconstructions toward low frequencies; (4) divergent self-consistency arising from a conflict with EMA training; and (5) curvy flow trajectories that impede convergence. To address these challenges, we introduce iSM, a unified training framework that systematically resolves each limitation. Our framework is built on four key improvements: Intrinsic Guidance provides explicit, dynamic control over guidance strength, resolving both compounding guidance and inflexibility. A Multi-Level Wavelet Loss mitigates frequency bias to restore high-frequency details. Scaling Optimal Transport (sOT) reduces training variance and learns straighter, more stable generative paths. Finally, a Twin EMA strategy reconciles training stability with self-consistency. Extensive experiments on ImageNet 256 x 256 demonstrate that our approach yields substantial FID improvements over baseline shortcut models across one-step, few-step, and multi-step generation, making shortcut models a viable and competitive class of generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。