评测大模型生成科学图像能力,发现复杂任务普遍表现不佳。
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
- 构建ScImage基准,从空间、数值、属性三方面评估科学图像生成
- GPT-4o在简单提示下表现尚可,复杂组合任务普遍失败
- 支持多语言输入,适合研究科学图文生成的学者参考
多模态大语言模型在文本到图像生成上表现优异,但在科学图像生成这一推动科研进展的关键应用中仍缺乏研究。本文提出ScImage基准,评估大模型在科学图像生成中的多维度理解能力,包括空间、数值和属性及其组合,聚焦科学对象(如矩形、圆形)间的关联。评估涵盖GPT-4o、Llama、AutomaTikZ、Dall-E和StableDiffusion五种模型,采用代码输出(Python、TikZ)与直接栅格图像生成两种模式,并测试英语、德语、波斯语和中文四种语言。由11位科学家基于正确性、相关性和科学准确性三个标准进行评估,结果显示:尽管GPT-4o在仅涉及单一维度(如空间或数值)的简单提示下生成质量尚可,但所有模型在复杂组合提示下均面临显著挑战。
原文摘要 · Abstract (English)
Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images--a critical application for accelerating scientific progress--remains underexplored. In this work, we address this gap by introducing ScImage, a benchmark designed to evaluate the multimodal capabilities of LLMs in generating scientific images from textual descriptions. ScImage assesses three key dimensions of understanding: spatial, numeric, and attribute comprehension, as well as their combinations, focusing on the relationships between scientific objects (e.g., squares, circles). We evaluate five models, GPT-4o, Llama, AutomaTikZ, Dall-E, and StableDiffusion, using two modes of output generation: code-based outputs (Python, TikZ) and direct raster image generation. Additionally, we examine four different input languages: English, German, Farsi, and Chinese. Our evaluation, conducted with 11 scientists across three criteria (correctness, relevance, and scientific accuracy), reveals that while GPT-4o produces outputs of decent quality for simpler prompts involving individual dimensions such as spatial, numeric, or attribute understanding in isolation, all models face challenges in this task, especially for more complex prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。