arXiv:2603.11414cs.CLcond-mat.mtrl-sci2026-03

评测大模型读图解材料科学题的能力,发现它们仍依赖记忆而非真看图。

MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models

  • 专为材料科学图像理解设计,用相图、应力应变曲线等真实图表考大模型
  • 137道大学级题目,多数模型仍靠记知识答对,难准确读图取数
  • 揭示大模型在数字精度和图像推理上的短板,适合研究多模态AI的学者

我们提出MaterialFigBench,一个用于评估多模态大语言模型(LLMs)解决大学水平材料科学问题能力的基准数据集,这些问题需准确解读图表才能得出正确答案。与依赖文本的现有基准不同,MaterialFigBench聚焦于相图、应力-应变曲线、阿伦尼乌斯图、衍射图谱和显微结构示意图等不可或缺的图像。数据集包含137道改编自标准教材的自由回答题,覆盖晶体结构、力学性能、扩散、相图、相变及材料电子性质等多个主题。为应对从图像中读取数值时的固有歧义,关键问题提供专家定义的答案范围。我们评估了包括ChatGPT和GPT系列在内的多个前沿多模态大模型,分析其在不同题型和版本中的表现。结果显示,尽管模型更新后整体准确率有所提升,当前模型仍缺乏真正的视觉理解与定量解析能力;许多正确答案源于记忆而非图像阅读。该基准揭示了视觉推理、数值精度和有效数字处理方面的持续缺陷,同时识别出部分性能已改善的问题类型。MaterialFigBench为推进材料科学领域多模态推理能力提供了系统化、领域专用的基础,有助于指导未来具备更强图表理解能力的大模型研发。

原文摘要 · Abstract (English)

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of figures. Unlike existing benchmarks that primarily rely on textual representations, MaterialFigBench focuses on problems in which figures such as phase diagrams, stress-strain curves, Arrhenius plots, diffraction patterns, and microstructural schematics are indispensable for deriving correct answers. The dataset consists of 137 free-response problems adapted from standard materials science textbooks, covering a broad range of topics including crystal structures, mechanical properties, diffusion, phase diagrams, phase transformations, and electronic properties of materials. To address unavoidable ambiguity in reading numerical values from images, expert-defined answer ranges are provided where appropriate. We evaluate several state-of-the-art multimodal LLMs, including ChatGPT and GPT models accessed via OpenAI APIs, and analyze their performance across problem categories and model versions. The results reveal that, although overall accuracy improves with model updates, current LLMs still struggle with genuine visual understanding and quantitative interpretation of materials science figures. In many cases, correct answers are obtained by relying on memorized domain knowledge rather than by reading the provided images. MaterialFigBench highlights persistent weaknesses in visual reasoning, numerical precision, and significant-digit handling, while also identifying problem types where performance has improved. This benchmark provides a systematic and domain-specific foundation for advancing multimodal reasoning capabilities in materials science and for guiding the development of future LLMs with stronger figure-based understanding.

多模态模型材料科学图像理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。