统一10余种单步扩散蒸馏方法,理论指导生成效果突破极限。
Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
- 基于f-散度扩展理论,设计可计算的训练损失函数。
- 在CIFAR10上达1.46(无条件)和1.38(有条件)新低FID值。
- 适用于文本到3D生成,质量与多样性优于现有方法。
本文将超过10种现有的单步扩散蒸馏方法(如Diff-Instruct、DMD、SIM、SiD、f-distill等)统一于一个理论驱动的框架——Uni-Instruct中。该框架基于我们提出的f-散度族扩散展开理论,解决了原始展开f-散度不可计算的问题,提出等价且可计算的损失函数,有效训练单步扩散模型以最小化展开f-散度。Uni-Instruct不仅从高层视角理解现有方法,还实现当前最优的生成性能。在CIFAR10上,无条件生成取得1.46的弗雷切特初始距离(FID),条件生成达1.38。在ImageNet-64×64上,单步生成FID达1.02,显著优于79步教师模型(2.35),提升1.33。该方法也成功拓展至文本到3D生成,在生成质量和多样性上略胜于SDS和VSD。
原文摘要 · Abstract (English)
In this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, $f$-distill, etc, inside a theory-driven framework which we name the \textbf{\emph{Uni-Instruct}}. Uni-Instruct is motivated by our proposed diffusion expansion theory of the $f$-divergence family. Then we introduce key theories that overcome the intractability issue of the original expanded $f$-divergence, resulting in an equivalent yet tractable loss that effectively trains one-step diffusion models by minimizing the expanded $f$-divergence family. The novel unification introduced by Uni-Instruct not only offers new theoretical contributions that help understand existing approaches from a high-level perspective but also leads to state-of-the-art one-step diffusion generation performances. On the CIFAR10 generation benchmark, Uni-Instruct achieves record-breaking Frechet Inception Distance (FID) values of \textbf{\emph{1.46}} for unconditional generation and \textbf{\emph{1.38}} for conditional generation. On the ImageNet-$64\times 64$ generation benchmark, Uni-Instruct achieves a new SoTA one-step generation FID of \textbf{\emph{1.02}}, which outperforms its 79-step teacher diffusion with a significant improvement margin of 1.33 (1.02 vs 2.35). We also apply Uni-Instruct on broader tasks like text-to-3D generation. For text-to-3D generation, Uni-Instruct gives decent results, which slightly outperforms previous methods, such as SDS and VSD, in terms of both generation quality and diversity. Both the solid theoretical and empirical contributions of Uni-Instruct will potentially help future studies on one-step diffusion distillation and knowledge transferring of diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。