arXiv:2410.22775cs.CV2024-10被引 4

对比扩散与自回归模型在复杂图文生成中的表现,发现扩散模型更优。

Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models

  • 用T2I-CompBench基准测试新旧模型的组合生成能力
  • FLUX扩散模型表现接近DALL-E3,LlamaGen自回归模型仍落后
  • 适合关注图文生成细节还原的研究者和开发者

文本到图像(T2I)生成模型如Stable Diffusion和DALL-E在生成高质量、逼真自然图像方面表现出色,但在捕捉输入提示中实体、属性和空间关系等细节时仍存在不足,尤其在新颖或复杂的组合场景下容易出现组合生成失败。近期推出的开源扩散模型FLUX展现出强大的图像生成能力,而自回归模型LlamaGen则宣称在视觉质量上可与扩散模型竞争。本研究使用T2I-CompBench基准评估这些新模型的组合生成能力,结果表明:在相同模型规模和推理时间条件下,未经优化的自回归模型LlamaGen尚未达到最先进的扩散模型水平;而开源扩散模型FLUX的表现已接近闭源领先模型DALL-E3。

原文摘要 · Abstract (English)

Text-to-image (T2I) generative models, such as Stable Diffusion and DALL-E, have shown remarkable proficiency in producing high-quality, realistic, and natural images from textual descriptions. However, these models sometimes fail to accurately capture all the details specified in the input prompts, particularly concerning entities, attributes, and spatial relationships. This issue becomes more pronounced when the prompt contains novel or complex compositions, leading to what are known as compositional generation failure modes. Recently, a new open-source diffusion-based T2I model, FLUX, has been introduced, demonstrating strong performance in high-quality image generation. Additionally, autoregressive T2I models like LlamaGen have claimed competitive visual quality performance compared to diffusion-based models. In this study, we evaluate the compositional generation capabilities of these newly introduced models against established models using the T2I-CompBench benchmark. Our findings reveal that LlamaGen, as a vanilla autoregressive model, is not yet on par with state-of-the-art diffusion models for compositional generation tasks under the same criteria, such as model size and inference time. On the other hand, the open-source diffusion-based model FLUX exhibits compositional generation capabilities comparable to the state-of-the-art closed-source model DALL-E3.

图文生成扩散模型自回归评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。