公平对比单步与多步图像生成模型,揭示引导强度的隐藏代价
Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models
- 统一采样步数与无分类器引导设置,实现跨模型公平比较
- 发现少量步骤下引导增强会提升FID但降低图文对齐与视觉质量
- 提出基于CLIP和人类偏好评分的标准化评估指标,揭示生成语义一致性
当前最先进的文生图模型生成质量高,但推理成本昂贵,需多步微分方程或去噪过程。原生单步模型通过一步将噪声映射为图像以降低开销,但与多步系统比较困难,因研究常使用不匹配的采样步数和不同的无分类器引导(CFG)设置,而CFG可使FID、Inception Score和基于CLIP的对齐度朝相反方向变化。此外,单步模型在多步推理下的扩展性尚不明确,且对标签条件生成器的分布外评估仍有限,仅限ImageNet。为此,我们在ImageNet验证集、ImageNetV2及新构建的与ImageNet标签一致的分布外数据集reLAIONet上,对八种模型(包括单步流模型MeanFlow、Improved MeanFlow、SoFlow,多步基线RAE、Scale-RAE,以及成熟系统SiT、Stable Diffusion 3.5、FLUX.1)进行类别条件协议下的基准测试。采用FID、Inception Score、CLIP Score和Pick Score评估,结果显示:在少步场景中,以FID为导向的模型优化与引导选择可能产生误导,引导调整虽改善FID,却损害图文对齐与人类偏好信号,反而降低视觉质量。为明确这些权衡,我们引入基于CLIP和Pick Score的尺度化版本的FID(csFID、psFID)与Inception Score(csIS、psIS),作为语义对齐生成的诊断工具。
原文摘要 · Abstract (English)
State-of-the-art text-to-image models produce high-quality images, but inference remains expensive as generation requires several sequential ODE or denoising steps. Native one-step models aim to reduce this cost by mapping noise to an image in a single step, yet fair comparisons to multi-step systems are difficult because studies use mismatched sampling steps and different classifier-free guidance (CFG) settings, where CFG can shift FID, Inception Score, and CLIP-based alignment in opposing directions. It is also unclear how well one-step models scale to multi-step inference, and there is limited standardized out-of-distribution evaluation for label-ID-conditioned generators beyond ImageNet. To address this, we benchmark eight models spanning one-step flows (MeanFlow, Improved MeanFlow, SoFlow), multi-step baselines (RAE, Scale-RAE), and established systems (SiT, Stable Diffusion 3.5, FLUX.1) under a class-conditional protocol on ImageNet validation, ImageNetV2, and reLAIONet, our new proofread out-of-distribution dataset aligned to ImageNet label IDs. Using FID, Inception Score, CLIP Score, and Pick Score, we show that FID-focused model development and CFG selection can be misleading in few-step regimes, where guidance changes can improve FID while degrading text-image alignment and human preference signals, worsening visual quality. To make these tradeoffs explicit, we introduce CLIP-scaled and PickScore-scaled variants of FID (csFID, psFID) and Inception Score (csIS, psIS) to serve as a diagnostic for semantically aligned image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。