arXiv:2503.20240cs.CV2025-03被引 8

用预训练模型的无条件噪声提升微调扩散模型的生成质量

Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models

  • 用基线模型的无条件噪声替代微调模型的无条件预测
  • 在图像与视频生成任务中显著提升条件生成质量
  • 无需修改训练流程,适合所有基于CFG的微调模型

Classifier-Free Guidance(CFG)是训练条件扩散模型的核心技术。通常做法是使用单一网络同时学习条件与无条件噪声预测,并以低丢弃率处理条件输入。然而我们发现,训练中无条件噪声学习带宽有限,导致无条件先验质量差,进而严重损害条件生成效果。受多数CFG模型通过微调具备优良无条件生成能力的基线模型启发,我们首次证明:仅用基线模型的无条件噪声替换微调模型中的无条件预测,即可显著提升条件生成性能。进一步实验表明,甚至可使用不同于微调模型的其他扩散模型提供无条件噪声。我们在多种基于CFG的条件生成模型上验证了该方法,涵盖图像与视频生成任务,包括Zero-1-to-3、Versatile Diffusion、DiT、DynamiCrafter和InstructPix2Pix。

原文摘要 · Abstract (English)

Classifier-Free Guidance (CFG) is a fundamental technique in training conditional diffusion models. The common practice for CFG-based training is to use a single network to learn both conditional and unconditional noise prediction, with a small dropout rate for conditioning. However, we observe that the joint learning of unconditional noise with limited bandwidth in training results in poor priors for the unconditional case. More importantly, these poor unconditional noise predictions become a serious reason for degrading the quality of conditional generation. Inspired by the fact that most CFG-based conditional models are trained by fine-tuning a base model with better unconditional generation, we first show that simply replacing the unconditional noise in CFG with that predicted by the base model can significantly improve conditional generation. Furthermore, we show that a diffusion model other than the one the fine-tuned model was trained on can be used for unconditional noise replacement. We experimentally verify our claim with a range of CFG-based conditional models for both image and video generation, including Zero-1-to-3, Versatile Diffusion, DiT, DynamiCrafter, and InstructPix2Pix.

扩散模型条件生成微调优化无条件先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。