arXiv:2502.10574cs.CV2025-02被引 19

动态调节生成过程中的引导强度,兼顾图文匹配与图像质量。

Classifier-free Guidance with Adaptive Scaling

  • 基于梯度自适应归一化稳定引导效果。
  • 采用时间相关曲线动态调整引导力度,提升生成一致性。
  • 适合追求高质量且精准对齐文本的图像生成用户。

Classifier-free guidance (CFG) 是现代文本驱动扩散模型中的关键机制。实践中,引导强度控制存在图像质量与文本匹配度之间的权衡:强引导使图像完美契合提示但降低质量,弱引导则提升质量却偏离提示。本文提出 $β$-CFG(基于 $β$-分布的自适应缩放),通过梯度驱动的自适应归一化稳定引导效果,并利用一族随时间变化的单模态 $β$-分布曲线,在扩散去噪过程中动态调节文本匹配与样本质量的平衡。实验表明,该方法在保持与参考 CFG 相当的 CLIP 语义相似度的同时,获得了更优的 FID 分数。

原文摘要 · Abstract (English)

Classifier-free guidance (CFG) is an essential mechanism in contemporary text-driven diffusion models. In practice, in controlling the impact of guidance we can see the trade-off between the quality of the generated images and correspondence to the prompt. When we use strong guidance, generated images fit the conditioned text perfectly but at the cost of their quality. Dually, we can use small guidance to generate high-quality results, but the generated images do not suit our prompt. In this paper, we present $β$-CFG ($β$-adaptive scaling in Classifier-Free Guidance), which controls the impact of guidance during generation to solve the above trade-off. First, $β$-CFG stabilizes the effects of guiding by gradient-based adaptive normalization. Second, $β$-CFG uses the family of single-modal ($β$-distribution), time-dependent curves to dynamically adapt the trade-off between prompt matching and the quality of samples during the diffusion denoising process. Our model obtained better FID scores, maintaining the text-to-image CLIP similarity scores at a level similar to that of the reference CFG.

扩散模型图像生成引导机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。