arXiv:2509.09307cs.CVcs.AI2025-09EMNLP被引 10

首个材料表征图像理解基准,揭示大模型在真实科研场景下的能力短板。

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

  • 构建涵盖21项任务的多模态评测基准MatCha
  • 模型在高阶专业问题上表现远逊于人类专家
  • 适合材料科学与智能科研代理研究者参考

材料表征是获取材料信息的基础,揭示加工-微观结构-性能关系,指导材料设计与优化。尽管多模态大语言模型(MLLMs)在材料科学的生成与预测任务中展现潜力,但其对真实表征图像数据的理解能力仍缺乏探索。为此,我们提出MatCha——首个材料表征图像理解基准,包含1500个需专家级领域知识的问题。该基准覆盖材料研究四个关键阶段,包含21项不同任务,均反映材料科学家面临的实际挑战。对先进MLLMs在MatCha上的评估显示,其性能显著落后于人类专家。这些模型在需要高级专业知识和复杂视觉感知的问题上表现下降,简单少样本及思维链提示难以缓解此局限。结果表明现有MLLMs在真实材料表征场景中适应性有限。我们希望MatCha能推动新材料发现与自主科学代理等领域的研究。MatCha已开源:https://github.com/FreedomIntelligence/MatCha。

原文摘要 · Abstract (English)

Materials characterization is fundamental to acquiring materials information, revealing the processing-microstructure-property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have recently shown promise in generative and predictive tasks within materials science, their capacity to understand real-world characterization imaging data remains underexplored. To bridge this gap, we present MatCha, the first benchmark for materials characterization image understanding, comprising 1,500 questions that demand expert-level domain expertise. MatCha encompasses four key stages of materials research comprising 21 distinct tasks, each designed to reflect authentic challenges faced by materials scientists. Our evaluation of state-of-the-art MLLMs on MatCha reveals a significant performance gap compared to human experts. These models exhibit degradation when addressing questions requiring higher-level expertise and sophisticated visual perception. Simple few-shot and chain-of-thought prompting struggle to alleviate these limitations. These findings highlight that existing MLLMs still exhibit limited adaptability to real-world materials characterization scenarios. We hope MatCha will facilitate future research in areas such as new material discovery and autonomous scientific agents. MatCha is available at https://github.com/FreedomIntelligence/MatCha.

材料科学多模态模型图像理解基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。