重新评估无需训练的图像生成增强方法,发现引导技术效果有限且不稳定。
Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models
- 在固定协议下对比八种无训练增强方法,基于组合对齐性评估
- 无方法在所有指标上超越原始引导,多数提升在误差范围内
- 注意力扰动在部分模型有效,但易导致其他模型性能下降
推理阶段的质量增强方法是无需昂贵重训练即可提升扩散模型性能的有效手段。本文研究一类源于无分类器引导(CFG)的训练无关技术,这些方法最初针对旧版U-Net扩散模型提出,并仅在孤立的图像质量指标上验证,未考虑生成图像与文本提示之间的组合一致性或语义对应关系。我们在两个开源权重的修正流变换器模型上,采用统一评估协议和三个组合对齐基准重新评估了其中八种方法。结果表明,没有任何方法在所有评测标准上持续优于原始CFG。APG虽取得若干名义最优成绩,但其增益常在评估不确定性范围内。注意力扰动方法在SD3.5 Medium上表现较好,但在FLUX.2 [klein] 4B Base上更频繁出现退化。相比之下,CFG仍是成本较低且表现稳健的基线。
原文摘要 · Abstract (English)
Inference-time quality-enhancement methods are an effective and widely adopted means of improving diffusion models without expensive retraining. We study a family of training-free techniques conceptually rooted in Classifier-Free Guidance (CFG), most of which were originally proposed on older U-Net diffusion models and validated using metrics that assess image quality in isolation, without accounting for compositional alignment or semantic correspondence between the generated image and its associated text prompt. We re-evaluate eight such methods on two open-weight rectified-flow transformers under a fixed per-model protocol and three compositional-alignment benchmarks. No method consistently improves on CFG across the measured criteria. APG obtains several nominal best scores, but the corresponding gains often remain within the estimated evaluation uncertainty. Attention-perturbation methods provide isolated gains on SD3.5 Medium and more frequent degradations on FLUX.2 [klein] 4B Base, while CFG remains a competitive lower-cost baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。