arXiv:2601.04498cs.LGcs.CV2026-01ACL被引 12

首个评估图文生成可靠性的基准,发现模型常犯数据错误。

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

  • 拆解可靠性为10类原子问题,用多模态大模型自动验证
  • 顶级模型问答准确率90%但整体图文正确率仅49%
  • 数据完整性等维度是普遍短板,适合关注生成可信度的研究者

信息图是结合数据可视化、文本和图像元素的复合视觉作品。尽管最近的文本到图像(T2I)模型能生成美观图像,但其在生成信息图方面的可靠性尚不明确。生成的信息图可能表面看似正确,实则存在数据编码失真或文本错误等隐蔽问题。本文提出IGENBENCH,首个评估文本到信息图生成可靠性的基准,包含600个精心设计的测试用例,覆盖30种信息图类型。我们设计了自动化评估框架,将可靠性验证分解为基于10类问题的二元判断,并利用多模态大语言模型(MLLMs)进行每项问题的验证,获得问题级准确率(Q-ACC)和信息图级准确率(I-ACC)。我们对10个最先进的T2I模型进行了全面评估。系统分析揭示关键洞见:(i) 三层次性能层级,最优模型达到Q-ACC 0.90,但I-ACC仅为0.49;(ii) 数据相关维度成为普遍瓶颈(如数据完整性:0.21);(iii) 所有模型均难以实现端到端的完全正确。IGENBENCH已开放获取:https://igen-bench.vercel.app/。

原文摘要 · Abstract (English)

Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their reliability in generating infographics remains unclear. Generated infographics may appear correct at first glance but contain easily overlooked issues, such as distorted data encoding or incorrect textual content. We present IGENBENCH, the first benchmark for evaluating the reliability of text-to-infographic generation, comprising 600 curated test cases spanning 30 infographic types. We design an automated evaluation framework that decomposes reliability verification into atomic yes/no questions based on a taxonomy of 10 question types. We employ multimodal large language models (MLLMs) to verify each question, yielding question-level accuracy (Q-ACC) and infographic-level accuracy (I-ACC). We comprehensively evaluate 10 state-of-the-art T2I models on IGENBENCH. Our systematic analysis reveals key insights for future model development: (i) a three-tier performance hierarchy with the top model achieving Q-ACC of 0.90 but I-ACC of only 0.49; (ii) data-related dimensions emerging as universal bottlenecks (e.g., Data Completeness: 0.21); and (iii) the challenge of achieving end-to-end correctness across all models. We release IGENBENCH at https://igen-bench.vercel.app/.

图文生成评估基准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。