动态调整生成引导尺度,提升文本到图像的生成质量与准确性
Navigating with Annealing Guidance Scale in Diffusion Space
- 根据条件噪声信号实时调整引导尺度,避免固定参数的局限性
- 在多个数据集上显著提升图像质量和文本对齐度,尤其在高引导值时表现更优
- 无需额外计算开销,可直接替代现有引导方法,适合快速部署
去噪扩散模型在文本条件图像生成中表现优异,但其效果高度依赖采样过程中的引导策略。分类器无关引导(CFG)通过设定引导尺度来平衡图像质量和文本对齐,然而该尺度的选择对收敛结果有显著影响。本文提出一种基于退火的引导调度器,根据条件噪声信号动态调整引导尺度。通过学习调度策略,有效缓解了传统CFG的敏感行为。实验表明,该方法显著提升了图像质量与文本对齐程度,且无需额外激活或内存开销,可无缝替代现有分类器无关引导,实现更优的文本对齐与图像质量权衡。
原文摘要 · Abstract (English)
Denoising diffusion models excel at generating high-quality images conditioned on text prompts, yet their effectiveness heavily relies on careful guidance during the sampling process. Classifier-Free Guidance (CFG) provides a widely used mechanism for steering generation by setting the guidance scale, which balances image quality and prompt alignment. However, the choice of the guidance scale has a critical impact on the convergence toward a visually appealing and prompt-adherent image. In this work, we propose an annealing guidance scheduler which dynamically adjusts the guidance scale over time based on the conditional noisy signal. By learning a scheduling policy, our method addresses the temperamental behavior of CFG. Empirical results demonstrate that our guidance scheduler significantly enhances image quality and alignment with the text prompt, advancing the performance of text-to-image generation. Notably, our novel scheduler requires no additional activations or memory consumption, and can seamlessly replace the common classifier-free guidance, offering an improved trade-off between prompt alignment and quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。