arXiv:2606.05949cs.CV2026-06

首个针对自然科学插图生成的细粒度评测基准,揭示当前模型在准确性和逻辑性上的短板。

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

论文配图:Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models
图 1 · 摘自论文原文
  • 构建跨学科、多布局的高质量科学插图数据集,支持细粒度元素标注。
  • 评估发现顶尖闭源模型仍存在文字渲染差、推理能力弱、生成与精确难兼顾等问题。
  • 适合关注科学图像生成质量、模型可解释性及可信度的研究者使用。

科学插图是自然科学研究中传达复杂概念和过程的重要工具。随着文本到图像(T2I)模型能力提升,其在科学插图生成中的应用日益增多。然而现有评测多停留在整体层面,忽视细粒度要素,且对科学推理能力和输出简洁性的量化不足。本文提出FEPBench,一个基于多学科、多布局高质量科学插图构建的基准,借助多模态大模型(MLLMs)与人工专家,提供细粒度原子级标注,并从指令忠实性、推理丰富性、语义精确性三个维度系统评估T2I模型。评估进一步分解至视觉、文本、关系与布局等元素。结果表明,即使是最先进的闭源模型(如GPT Image 2和Nano Banana Pro),仍存在文字渲染瓶颈、推理丰富性不足、生成丰富性与精确性难以平衡等问题。研究为提升和部署T2I模型于科学插图生成提供了实践指导。相关数据、标注与代码将公开发布。

原文摘要 · Abstract (English)

Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts and processes. As Text-to-Image (T2I) models become increasingly capable, researchers have started to use them for scientific illustration generation. However, existing benchmarks often assess outputs at a holistic level, overlooking fine-grained elements, while scientific reasoning ability and output conciseness remain under-quantified. We introduce FEPBench, a benchmark built from carefully selected high-quality scientific illustrations across multiple disciplines and layout types. With the assistance of multimodal large language models (MLLMs) and human experts, we provide fine-grained atom set annotations and systematically evaluate T2I models along three dimensions: instruction faithfulness, reasoning enrichment, and semantic precision. Our evaluation further decomposes model performance across visual, textual, relation, and layout elements. Results show that even state-of-the-art (SOTA) closed-source models, such as GPT Image 2 and Nano Banana Pro, still suffer from text-rendering bottlenecks, limited reasoning enrichment, and difficulty balancing generation richness with precision. These findings provide practical guidance for improving and deploying T2I models in scientific illustration generation. Benchmark data, atom set annotations, and evaluation code will be released by us.

图像生成科学可视化评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。