现有图像生成评估方法被引导强度误导,真实性能可能被高估。
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
- 发现人类偏好模型对高引导强度有严重偏倚
- 提出新评估框架GA-Eval,消除引导强度干扰
- 揭示多数新方法在真实场景中效果不如标准引导
无分类器引导(CFG)已推动扩散模型在文本到图像生成中取得显著进展。然而,近期涌现的多种引导方法是否真有实质性提升?本文揭示一个关键评估陷阱:常见人类偏好模型对大引导强度存在强偏倚。仅提高CFG强度即可显著提升量化评分,即使图像质量严重受损(如过饱和、伪影)。为此,我们提出引导感知评估(GA-Eval)框架,通过有效校准引导强度,实现对当前引导方法与CFG的公平比较。受此启发,我们设计了超凡扩散引导(TDG),其在传统评估下表现优异,但实际效果不佳。在广泛实验中,我们对比了八种最新引导方法在传统与GA-Eval框架下的表现,发现仅提升CFG强度即可媲美多数方法,且所有方法在标准CFG下均出现胜率严重下降。本工作强烈呼吁学界重新思考该领域的评估范式与未来方向。
原文摘要 · Abstract (English)
Classifier-free guidance (CFG) has helped diffusion models achieve great conditional generation in various fields. Recently, more diffusion guidance methods have emerged with improved generation quality and human preference. However, can these emerging diffusion guidance methods really achieve solid and significant improvements? In this paper, we rethink recent progress on diffusion guidance. Our work mainly consists of four contributions. First, we reveal a critical evaluation pitfall that common human preference models exhibit a strong bias towards large guidance scales. Simply increasing the CFG scale can easily improve quantitative evaluation scores due to strong semantic alignment, even if image quality is severely damaged (e.g., oversaturation and artifacts). Second, we introduce a novel guidance-aware evaluation (GA-Eval) framework that employs effective guidance scale calibration to enable fair comparison between current guidance methods and CFG by identifying the effects orthogonal and parallel to CFG effects. Third, motivated by the evaluation pitfall, we design Transcendent Diffusion Guidance (TDG) method that can significantly improve human preference scores in the conventional evaluation framework but actually does not work in practice. Fourth, in extensive experiments, we empirically evaluate recent eight diffusion guidance methods within the conventional evaluation framework and the proposed GA-Eval framework. Notably, simply increasing the CFG scales can compete with most studied diffusion guidance methods, while all methods suffer severely from winning rate degradation over standard CFG. Our work would strongly motivate the community to rethink the evaluation paradigm and future directions of this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。