arXiv:2506.08351cs.CVcs.AI2025-06被引 2

只在前几步用引导,生成速度提升20%-30%。

How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models

  • 只在去噪前期使用分类器无关引导,减少计算量。
  • 在不同模型和步骤下,图像质量与图文对齐保持不变。
  • 方法简单通用,适合各类扩散模型加速推理。

随着文本到图像生成扩散模型的快速发展,分类器无关引导已成为主流条件生成方法。然而,该方法相比无条件生成需多一倍前向计算,显著增加成本。尽管已有研究提出自适应引导概念,但缺乏充分分析与实证,难以推广至通用扩散模型。本文提出一种新视角下的自适应引导策略 Step AG,简单且普适。评估聚焦图像质量与图文对齐,结果表明:仅在前若干去噪步骤中应用分类器无关引导即可生成高质量、良好对齐的图像,平均提速20%至30%。该优势在不同推理步数及多种模型(包括视频生成模型)中均一致,凸显方法优越性。

原文摘要 · Abstract (English)

With the rapid development of text-to-vision generation diffusion models, classifier-free guidance has emerged as the most prevalent method for conditioning. However, this approach inherently requires twice as many steps for model forwarding compared to unconditional generation, resulting in significantly higher costs. While previous study has introduced the concept of adaptive guidance, it lacks solid analysis and empirical results, making previous method unable to be applied to general diffusion models. In this work, we present another perspective of applying adaptive guidance and propose Step AG, which is a simple, universally applicable adaptive guidance strategy. Our evaluations focus on both image quality and image-text alignment. whose results indicate that restricting classifier-free guidance to the first several denoising steps is sufficient for generating high-quality, well-conditioned images, achieving an average speedup of 20% to 30%. Such improvement is consistent across different settings such as inference steps, and various models including video generation models, highlighting the superiority of our method.

扩散模型文本生成图像加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。