构建材料相图理解基准,评估AI是否真懂科学图像的深层机制。
MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding

- 基于3681篇论文筛选200组高质量图文对,覆盖189种材料系统
- 当前视觉语言模型仅能识别表面特征,无法理解热力学机理
- 适合需要深度科学推理能力的科研人员和多模态AI开发者
材料相图是材料科学的核心知识表征,包含温度、成分、相稳定性及相变路径等信息,其完整理解需依赖热力学机制分析与科学推理。尽管视觉语言模型(VLMs)在科学图像理解方面展现出潜力,但对其在需深层机制解读的复杂图像上的系统性评估仍有限,而相图为此提供了极具挑战性的测试平台。本文提出MatPhaseBench,一个高质量、高可靠性的复杂科学图像理解基准,聚焦材料相图。该基准源自3681篇经典材料学期刊论文,从中精选200组高质量图-文配对,涵盖189种材料体系与70种元素。其三大特点:(1)面向复杂科学图像理解——超越简单目标识别,聚焦需深度认知的开放任务;(2)全面的图文语义对齐——文献挖掘与匹配过程中完整保留图像相关语义信息;(3)高质量人工监督文本获取——所有描述均经严格人工验证。实验结果表明,当前VLMs仍远低于专家水平:主要局限于表面视觉感知,缺乏基于热力学机制的深度推理,领域意识薄弱,且在复合或多重图场景中难以区分细微差异。总体而言,MatPhaseBench构成一项具有挑战性的研究级基准,为复杂科学图像理解、相图分析及可信多模态科学人工智能提供基础平台。
原文摘要 · Abstract (English)
Materials phase diagrams are a core knowledge representation in materials science, encoding temperature,composition, phase stability, and phase transformation pathways, with their full understanding requiring thermodynamic mechanism analysis and scientific reasoning. Although VLMs have shown promise in scientific image understanding, their systematic evaluation on such logically complex images demanding deep mechanistic interpretation remains limited, and phase diagrams provide a challenging testbed for this purpose. We introduce MatPhaseBench, a high-quality, high-reliability benchmark for complex scientific image understanding, focused on materials phase diagrams. MatPhaseBench is constructed from 3681 papers in classical materials science journals, from which 200 high-quality diagram-text pairs were selected, covering 189 material systems and 70 elements. The benchmark has three key features: (1)targeting complex scientific image understanding-it moves beyond simple objective tests to open-ended tasks requiring deep comprehension; (2)comprehensive image-text alignment-semantic information associated with images is fully preserved during literature mining and matching; (3) high-quality human-supervised text acquisition-all descriptions undergo strict manual validation. Experimental results show that current VLMs remain substantially behind expert-level understanding: they are largely limited to surface visual perception, lack deep reasoning grounded in thermodynamic mechanisms, have limited domain awareness and expert analytical experience, and perform poorly in distinguishing fine-grained differences in composite or multi-diagram settings. Overall, MatPhaseBench constitutes a challenging research-grade benchmark, providing a foundational platform for complex scientific image understanding, phase diagram analysis, and trustworthy multi-modal AI in science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。