提出动态调节生成引导强度的新方法,提升图像质量和多样性。
Feedback Guidance of Diffusion Models
- 根据生成质量自适应调整引导强度,避免固定参数带来的缺陷。
- 在ImageNet上优于传统无分类器引导,接近最优区间引导效果。
- 可自动识别复杂提示并增强引导,兼容现有引导策略。
尽管无分类器引导(CFG)已成为条件扩散模型提升样本保真度的标准方法,但其恒定的引导强度会损害多样性并引发记忆现象。本文提出反馈引导(FBG),通过状态依赖系数动态调节引导强度,依据生成需求自我校准。该方法基于条件分布线性受无条件分布干扰的假设,不同于CFG隐含的乘法关系。其核心在于利用自身对条件信号信息量的预测反馈,实现推理过程中的动态引导,挑战了将引导视为固定超参数的传统观点。在ImageNet 512x512数据集上,该方法显著优于无分类器引导,并在性能上与有限区间引导(LIG)相当,同时具备坚实的数学基础。在文生图任务中,结果显示其能自动为复杂提示施加更高引导强度,且可无缝集成现有引导机制如CFG或LIG。
原文摘要 · Abstract (English)
While Classifier-Free Guidance (CFG) has become standard for improving sample fidelity in conditional diffusion models, it can harm diversity and induce memorization by applying constant guidance regardless of whether a particular sample needs correction. We propose FeedBack Guidance (FBG), which uses a state-dependent coefficient to self-regulate guidance amounts based on need. Our approach is derived from first principles by assuming the learned conditional distribution is linearly corrupted by the unconditional distribution, contrasting with CFG's implicit multiplicative assumption. Our scheme relies on feedback of its own predictions about the conditional signal informativeness to adapt guidance dynamically during inference, challenging the view of guidance as a fixed hyperparameter. The approach is benchmarked on ImageNet512x512, where it significantly outperforms Classifier-Free Guidance and is competitive to Limited Interval Guidance (LIG) while benefitting from a strong mathematical framework. On Text-To-Image generation, we demonstrate that, as anticipated, our approach automatically applies higher guidance scales for complex prompts than for simpler ones and that it can be easily combined with existing guidance schemes such as CFG or LIG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。