在复杂图文提示下对比主流文生图模型表现,揭示真实差距。
Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

- 用48个高难度提示测试四大文生图系统,聚焦精确计数和空间约束。
- 谷歌Gemini 3 Pro Image以84.8分领先,领先第二名FLUX.2约2.5分。
- 顶尖模型主要错在物体数量错误和几何失真,落后者多出现文字乱码。
文生图模型通常在平均难度提示下评估,这掩盖了系统在涉及精确物体数量、多对象属性绑定、可读嵌入文本和明确空间约束等复合需求下的真实差距。本文评估了四个生产级文生图系统:Hunyuan 3.0、Gemini 3 Pro Image(“Nano Banana Pro”)、Black Forest Labs FLUX.2 和 Ideogram 3.0。评估基于 DataSeeds.AI 样本数据集(DSD)中通过自动化复杂度评分筛选出的48个最困难提示。每张生成图像均由独立评审员打分,使用 GPT-5.4-Pro 构建的原子化、加权、互斥且穷尽(MECE)评分标准,同时 Gemini 3.1 Pro Preview 独立判断各项标准是否满足。结果显示,Gemini 3 Pro Image 得分为 84.8/100,略胜于 FLUX.2 的 82.3/100;Ideogram 3.0 和 Hunyuan 3.0 分别得 65.7/100 和 63.3/100。故障分析表明,领先系统主要因物体计数错误和几何失真丢分,而落后系统更常出现文字乱码;Ideogram 3.0 还频繁遗漏请求元素。完整样本评分表、得分与故障标注可向作者索取。
原文摘要 · Abstract (English)
Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded text, and explicit spatial constraints. We evaluate four production text-to-image systems: Hunyuan 3.0, Gemini 3 Pro Image ("Nano Banana Pro"), Black Forest Labs FLUX.2, and Ideogram 3.0. The evaluation uses the 48 hardest prompts drawn from the DataSeeds.AI Sample Dataset (DSD), selected through an automated complexity-scoring pass over the full corpus. Every generated image is graded using an independent-judge rubric. GPT-5.4-Pro authors an atomic, weighted, mutually exclusive and collectively exhaustive (MECE) evaluation rubric, while Gemini 3.1 Pro Preview independently determines whether each criterion is satisfied. Gemini 3 Pro Image ranks first with a score of 84.8/100, narrowly ahead of FLUX.2 at 82.3/100. Ideogram 3.0 and Hunyuan 3.0 score 65.7/100 and 63.3/100, respectively. Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text. Ideogram 3.0 also frequently omits requested elements. Full per-sample rubrics, scores, and failure annotations are available from the authors upon request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。