arXiv:2505.22126cs.CVcs.AI2025-05被引 19

首个科学图示生成基准,评估模型将科研内容转为准确图示的能力

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

  • 构建涵盖13个学科的1120个科学图示样本,由专家与多模态模型共同筛选
  • 顶级模型GPT-4o-image在语义一致性和结构准确性上仍显著低于人类表现
  • 聚焦科学严谨性,适合关注科研可视化与智能作图的研究者参考

近年来,基于AI的图像生成技术发展迅速。早期扩散模型侧重感知质量,而GPT-4o-image等新型多模态模型则融合高层推理,提升语义理解与结构组织能力。科学插图生成体现了这一演进:不同于通用图像合成,其需准确解析技术内容,并将抽象概念转化为清晰、标准化的视觉表达,任务更具知识密集性且耗时费力,常需数小时人工操作与专用工具。实现可控、智能化自动化具有重要实用价值。然而,目前尚无针对该方向的评测基准。为此,我们提出SridBench,首个科学图表生成基准。其包含1,120个样本,源自13个自然科学与计算机科学领域的顶尖论文,通过人工专家与多模态大模型(MLLMs)协同收集。每项样本从语义保真度、结构准确性等六个维度评估。实验表明,即使顶级模型GPT-4o-image也显著落后于人类表现,常见问题包括文本/视觉清晰度不足与科学正确性偏差。结果凸显了亟需更强推理驱动的视觉生成能力。

原文摘要 · Abstract (English)

Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate high-level reasoning, improving semantic understanding and structural composition. Scientific illustration generation exemplifies this evolution: unlike general image synthesis, it demands accurate interpretation of technical content and transformation of abstract ideas into clear, standardized visuals. This task is significantly more knowledge-intensive and laborious, often requiring hours of manual work and specialized tools. Automating it in a controllable, intelligent manner would provide substantial practical value. Yet, no benchmark currently exists to evaluate AI on this front. To fill this gap, we introduce SridBench, the first benchmark for scientific figure generation. It comprises 1,120 instances curated from leading scientific papers across 13 natural and computer science disciplines, collected via human experts and MLLMs. Each sample is evaluated along six dimensions, including semantic fidelity and structural accuracy. Experimental results reveal that even top-tier models like GPT-4o-image lag behind human performance, with common issues in text/visual clarity and scientific correctness. These findings highlight the need for more advanced reasoning-driven visual generation capabilities.

科学绘图多模态生成评测基准图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。