arXiv:2603.27862cs.GRcs.AI2026-03被引 6

用3600个真实任务测试图像生成模型,揭示其在编辑和文字密集场景下的短板。

ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

  • 构建涵盖6类任务与6大领域的3600个条件集,支持细粒度人类标注。
  • 模型在编辑任务中表现差于生成任务,尤其在局部修改时出错率高。
  • 适合关注图像生成鲁棒性、评测方法改进的研究者与开发者。

扩散、自回归及混合模型的进展使文本到图像、编辑与参考引导构图等任务实现高质量图像合成。然而,现有基准仍存在局限:或仅聚焦孤立任务,或覆盖领域狭窄,或评分不透明、无法解释失败原因。本文提出 extbf{ImagenWorld},一个包含3.6K条件集的基准,涵盖六项核心任务(生成与编辑,单/多参考)和六个主题领域(艺术品、照片级图像、信息图表、文本图形、计算机图形、截图)。该基准依托20,000条细粒度人工标注,采用可解释的评估框架,对局部对象级与区域级错误进行标记,补充基于VLM的自动化指标。对14个模型的大规模评估显示:(1) 模型在编辑任务中表现普遍弱于生成任务,尤其在局部编辑中更易出错;(2) 在艺术与照片级场景中表现优异,但在符号化与文字密集领域如截图与信息图表中表现不佳;(3) 封闭源模型整体领先,而针对性数据训练(如Qwen-Image)可缩小文字密集场景中的差距;(4) 现代VLM指标最高达Kendall相关性0.79,接近人类排序,但难以提供细粒度错误归因。ImagenWorld既为严谨评测,也为诊断工具,推动图像生成的稳健发展。

原文摘要 · Abstract (English)

Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated tasks, cover only narrow domains, or provide opaque scores without explaining failure modes. We introduce \textbf{ImagenWorld}, a benchmark of 3.6K condition sets spanning six core tasks (generation and editing, with single or multiple references) and six topical domains (artworks, photorealistic images, information graphics, textual graphics, computer graphics, and screenshots). The benchmark is supported by 20K fine-grained human annotations and an explainable evaluation schema that tags localized object-level and segment-level errors, complementing automated VLM-based metrics. Our large-scale evaluation of 14 models yields several insights: (1) models typically struggle more in editing tasks than in generation tasks, especially in local edits. (2) models excel in artistic and photorealistic settings but struggle with symbolic and text-heavy domains such as screenshots and information graphics. (3) closed-source systems lead overall, while targeted data curation (e.g., Qwen-Image) narrows the gap in text-heavy cases. (4) modern VLM-based metrics achieve Kendall accuracies up to 0.79, approximating human ranking, but fall short of fine-grained, explainable error attribution. ImagenWorld provides both a rigorous benchmark and a diagnostic tool to advance robust image generation.

图像生成评测基准可解释性VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。