arXiv:2505.23982cs.AI2025-05被引 5

首个评估大模型材料科学推理能力的基准,涵盖1757道研究生级题目。

MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge

  • 构建包含1757道题目的双格式评测集,覆盖七大学科方向
  • 开源模型最高准确率60.5%,闭源模型达84.5%仍有显著差距
  • 适合材料科学与AI交叉研究者参考,推动领域专用模型发展

尽管大语言模型在材料科学领域取得进展,但缺乏对其领域知识与复杂推理能力的评估基准。为此,我们提出MSQA,一个包含1,757道研究生级材料科学问题的综合性评测基准,支持详细解释性回答和二元真假判断两种形式。该基准通过要求模型在结构-性能关系、合成工艺、计算建模等七个子领域进行多步推理,对大模型提出了更高挑战。实验对比了10个前沿大模型,发现闭源模型最高准确率达84.5%,而开源模型仅约60.5%,且领域专用模型常因过拟合与分布偏移表现不佳。MSQA是首个同时评估大模型事实知识与推理能力的基准,对推进高级材料科学中的大模型应用具有重要意义。

原文摘要 · Abstract (English)

Despite recent advances in large language models (LLMs) for materials science, there is a lack of benchmarks for evaluating their domain-specific knowledge and complex reasoning abilities. To bridge this gap, we introduce MSQA, a comprehensive evaluation benchmark of 1,757 graduate-level materials science questions in two formats: detailed explanatory responses and binary True/False assessments. MSQA distinctively challenges LLMs by requiring both precise factual knowledge and multi-step reasoning across seven materials science sub-fields, such as structure-property relationships, synthesis processes, and computational modeling. Through experiments with 10 state-of-the-art LLMs, we identify significant gaps in current LLM performance. While API-based proprietary LLMs achieve up to 84.5% accuracy, open-source (OSS) LLMs peak around 60.5%, and domain-specific LLMs often underperform significantly due to overfitting and distributional shifts. MSQA represents the first benchmark to jointly evaluate the factual and reasoning capabilities of LLMs crucial for LLMs in advanced materials science.

材料科学大模型评测多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。