首个评估科学图生成的基准,专测标签准确与图示规范性。
Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

- 构建32项跨10学科的科学图生成任务,含可机器验证的规范要求
- 在8类图上,专用模型在语义正确性和规范遵循上显著领先通用模型
- 提出四维评测体系,聚焦文本可读性、语义准确性等科学图核心质量
文本到图像及多模态生成模型正被用于生成机制图、实验设计示意图、概念框架和图形摘要等科学图。然而现有图像生成基准(如GenEval、T2I-CompBench、DPG-Bench)仅评估自然图像的构图、物体计数或写实度,未涵盖科学图的关键可用性指标:标签正确且清晰、实体及其关系忠实呈现、图示结构连贯、符合学科绘图惯例。我们提出SciDraw-Bench,包含32个结构化科学图生成任务,覆盖8种图类型和10个学科领域,每项任务均配以自然语言提示和机器可验证的规范说明,包括所需标签、关系、组件、惯例及负向约束。我们设计四维评估协议:文本保真度(基于OCR的标签召回率与字符错误率)、语义正确性(视觉语言模型依据规范判断)、结构质量与惯例遵循度,并引入元评估与初步人评一致性分析(人工评分验证正在进行中)。我们在代表性通用文本到图像模型上评估了领域专用系统SciDraw AI,同时规划了代码到图的基线。在全部8类图的预实验中,专用系统在各维度和图类型上均显著优于通用基线,尤其在语义正确性和惯例遵循方面差距最大;而文本保真度对所有系统仍是难点。
原文摘要 · Abstract (English)
Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts. Yet existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or photorealism. None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coherent diagrammatic structure, and adherence to disciplinary drawing conventions. We introduce SciDraw-Bench, a benchmark of 32 structured scientific-figure generation tasks spanning eight figure types and ten disciplines, where each task pairs a natural-language prompt with a machine-checkable specification of required labels, relations, components, conventions, and negative constraints. We propose a four-dimensional evaluation protocol: Text Fidelity (OCR-based label recall and character error rate), Semantic Correctness (vision-language-model judging against the specification), Structural Quality, and Convention Adherence, together with a meta-evaluation protocol and a preliminary inter-judge reliability analysis (human-rating validation is ongoing). We evaluate a domain-specific system, SciDraw AI, against representative general-purpose text-to-image models, and outline a code-to-figure baseline as a planned extension. In a pilot over all eight figure types, the domain-specific system substantially outperforms the general-purpose baselines on every dimension and figure type, with the largest gaps on semantic correctness and convention adherence; text fidelity remains the hardest dimension for all systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。