评测顶尖文生图模型在多场景下的生成能力,揭示其向通用AI工具演进的潜力。
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
- 构建跨五领域的综合评测框架,涵盖结构化输出与物理一致性等维度。
- FLUX.1和Ideogram2.0在结构化任务中表现领先,其他模型各有专长。
- 适合关注文生图模型能力边界与实际应用落地的研究者与开发者。
随着扩散模型的快速发展,文生图(T2I)模型在提示遵循与图像生成方面取得显著进展。新发布的FLUX.1、Ideogram2.0,以及Dall-E3、Stable Diffusion 3等模型在复杂任务中展现出卓越性能,引发对其迈向通用应用的思考。这些模型已扩展至可控生成、图像编辑、视频、音频、3D与运动生成,以及语义分割、深度估计等计算机视觉任务。然而,现有评估框架难以全面衡量其跨域表现。为此,我们提出IMAGINE-E评测体系,测试了六款主流模型:FLUX.1、Ideogram2.0、Midjourney、Dall-E3、Stable Diffusion 3和Jimeng。评估涵盖五大核心领域:结构化输出生成、真实性与物理一致性、特定领域生成、挑战性场景生成及多风格创作任务。结果揭示各模型优劣势,尤其凸显FLUX.1与Ideogram2.0在结构化与特定领域任务中的突出表现,表明T2I模型正朝着通用基础模型方向演进。本研究为理解当前状态与未来趋势提供关键洞察。评测脚本将开源于https://github.com/jylei16/Imagine-e。
原文摘要 · Abstract (English)
With the rapid development of diffusion models, text-to-image(T2I) models have made significant progress, showcasing impressive abilities in prompt following and image generation. Recently launched models such as FLUX.1 and Ideogram2.0, along with others like Dall-E3 and Stable Diffusion 3, have demonstrated exceptional performance across various complex tasks, raising questions about whether T2I models are moving towards general-purpose applicability. Beyond traditional image generation, these models exhibit capabilities across a range of fields, including controllable generation, image editing, video, audio, 3D, and motion generation, as well as computer vision tasks like semantic segmentation and depth estimation. However, current evaluation frameworks are insufficient to comprehensively assess these models' performance across expanding domains. To thoroughly evaluate these models, we developed the IMAGINE-E and tested six prominent models: FLUX.1, Ideogram2.0, Midjourney, Dall-E3, Stable Diffusion 3, and Jimeng. Our evaluation is divided into five key domains: structured output generation, realism, and physical consistency, specific domain generation, challenging scenario generation, and multi-style creation tasks. This comprehensive assessment highlights each model's strengths and limitations, particularly the outstanding performance of FLUX.1 and Ideogram2.0 in structured and specific domain tasks, underscoring the expanding applications and potential of T2I models as foundational AI tools. This study provides valuable insights into the current state and future trajectory of T2I models as they evolve towards general-purpose usability. Evaluation scripts will be released at https://github.com/jylei16/Imagine-e.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。