提出一种无需额外模型的扩散模型引导方法,同时提升图像质量、多样性与提示一致性。
Entropy Rectifying Guidance for Diffusion and Flow Models
- 基于扩散模型注意力机制在推理时的熵变化,设计简单有效的引导策略。
- 在文生图、类别条件生成等任务中显著提升图像质量与多样性,且保持提示一致性。
- 兼容其他先进引导方法,适用于有指导和无指导生成,适合追求高质量生成的用户。
引导技术广泛应用于扩散模型和流模型中,以提升分类条件和文本到图像生成等任务中的图像质量和输入一致性。其中,无分类器引导(CFG)是最常用的引导方法,但其在质量、多样性和一致性之间存在权衡:提升一项往往牺牲另一项。尽管近期工作已部分实现三者的解耦,但这些方法需额外(较弱)模型或每步采样增加前向传播次数。本文提出熵修正引导(ERG),一种基于当前最先进的扩散变压器架构在推理时注意力机制变化的简单有效引导方法,可同时提升图像质量、多样性与提示一致性。ERG比CFG等方法更通用,支持无条件采样。实验表明,ERG在文本到图像、类别条件生成及无条件图像生成等多种任务中均有显著提升,并能无缝结合如CADS和APG等最新引导方法,进一步优化生成效果。
原文摘要 · Abstract (English)
Guidance techniques are commonly used in diffusion and flow models to improve image quality and input consistency for conditional generative tasks such as class-conditional and text-to-image generation. In particular, classifier-free guidance (CFG) is the most widely adopted guidance technique. It results, however, in trade-offs across quality, diversity and consistency: improving some at the expense of others. While recent work has shown that it is possible to disentangle these factors to some extent, such methods come with an overhead of requiring an additional (weaker) model, or require more forward passes per sampling step. In this paper, we propose Entropy Rectifying Guidance (ERG), a simple and effective guidance method based on inference-time changes in the attention mechanism of state-of-the-art diffusion transformer architectures, which allows for simultaneous improvements over image quality, diversity and prompt consistency. ERG is more general than CFG and similar guidance techniques, as it extends to unconditional sampling. We show that ERG results in significant improvements in various tasks, including text-to-image, class-conditional and unconditional image generation. We also show that ERG can be seamlessly combined with other recent guidance methods such as CADS and APG, further improving generation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。