让一致性模型无需教师模型即可实现可调引导生成。
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
- 通过联合流分布学习,使预训练模型实现后置引导。
- 在CIFAR-10和ImageNet 64x64上降低FID,提升生成质量。
- 首次实现一致性模型无依赖式引导,适合图像生成研究者。
分类器无关引导(CFG)可调节扩散模型的保真度与多样性,但其应用受限于采样成本。一致性模型(CMs)虽能在一至几步内生成图像,现有引导方法需从独立的扩散模型教师中进行知识蒸馏,仅限于一致性蒸馏(CD)方法。本文提出联合流分布学习(JFDL),一种轻量级对齐方法,使预训练的CM具备引导能力。以预训练的CM作为常微分方程(ODE)求解器,通过正态性检验验证了无条件与条件分布速度场所隐含的方差爆炸噪声为高斯分布。实践中,JFDL为CM提供了熟悉的可调引导旋钮,生成图像特性与CFG相似。应用于原本仅支持条件采样的原始一致性训练(CT)CM,JFDL实现了引导生成,并在CIFAR-10与ImageNet 64x64数据集上降低FID。这是首次实现一致性模型无需扩散模型教师即可有效引导,填补了当前一致性模型方法的关键空白。
原文摘要 · Abstract (English)
Classifier-free Guidance (CFG) lets practitioners trade-off fidelity against diversity in Diffusion Models (DMs). The practicality of CFG is however hindered by DMs sampling cost. On the other hand, Consistency Models (CMs) generate images in one or a few steps, but existing guidance methods require knowledge distillation from a separate DM teacher, limiting CFG to Consistency Distillation (CD) methods. We propose Joint Flow Distribution Learning (JFDL), a lightweight alignment method enabling guidance in a pre-trained CM. With a pre-trained CM as an ordinary differential equation (ODE) solver, we verify with normality tests that the variance-exploding noise implied by the velocity fields from unconditional and conditional distributions is Gaussian. In practice, JFDL equips CMs with the familiar adjustable guidance knob, yielding guided images with similar characteristics to CFG. Applied to an original Consistency Trained (CT) CM that could only do conditional sampling, JFDL unlocks guided generation and reduces FID on both CIFAR-10 and ImageNet 64x64 datasets. This is the first time that CMs are able to receive effective guidance post-hoc without a DM teacher, thus, bridging a key gap in current methods for CMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。