首个材料科学推理评测基准,揭示大模型在专业领域的真实短板
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- 构建涵盖1340题的材料科学多层级评测集,含图文与分步解法
- 顶尖模型文本题准确率仅75.22%,图像题最高53.02%仍不理想
- 工具增强可高效提升非思考模型,自修正反而常导致错误
大型语言模型虽展现强大科学推理能力,但在材料科学领域的表现仍缺乏系统评估。为此,我们提出MatSciBench,一个涵盖1340道大学水平材料科学问题的综合性基准,覆盖6个主领域和31个子领域,并按推理长度分为三类难度。该基准包含946道题的详细参考解答,支持过程级错误分析,以及315道带图题目以评估多模态推理能力。我们在MatSciBench上评测了主流思维与非思维型LLM,并测试三种非思维模型的推理增强方法:基础链式思维提示、工具增强与自我修正。结果表明,当前模型在大学级材料科学推理中仍有明显局限:DeepSeek-R1在纯文本题上取得最高75.22%准确率,GPT-5在图文题上达53.02%最优。分析显示,工具增强能以较低代价有效提升多数非思维模型性能,而自我修正常无法带来可靠增益,甚至会将正确答案改为错误。进一步分析揭示,现有模型主要受限于领域知识缺失、计算错误、问题理解偏差及从科学图表中提取精确信息的能力不足。总体而言,MatSciBench为评估大模型在材料科学中的推理能力提供了清晰基准,有助于推动未来研究。
原文摘要 · Abstract (English)
Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To fill this gap, we introduce MatSciBench, a comprehensive college-level benchmark comprising 1340 problems that span the essential subdisciplines of materials science. MatSciBench features a structured and fine-grained taxonomy that categorizes materials science questions into 6 primary fields and 31 subfields, together with a three-tier difficulty classification based on the reasoning length needed to solve each problem. MatSciBench includes detailed reference solutions for 946 questions, supports process-level error analysis, and contains 315 questions with images for evaluating multimodal reasoning. We evaluate leading thinking and non-thinking LLMs on MatSciBench, and further test three reasoning methods for non-thinking models: basic chain-of-thought prompting, tool augmentation, and self-correction. The results show that current models still face clear limits in college-level materials science reasoning. DeepSeek-R1 achieves the highest score on text-only questions at 75.22% accuracy, and GPT-5 performs the best on questions with images at 53.02%. Our analysis shows that tool augmentation improves many non-thinking models in a token-efficient way, while self-correction often fails to provide reliable gains and can revise correct answers into incorrect ones. We further analyze performance across difficulty levels, reasoning efficiency, multimodal reasoning, and failure patterns, and find that current models are mainly limited by domain knowledge gaps, calculation errors, problem comprehension failures, and difficulty in extracting precise information from scientific figures. Overall, MatSciBench provides a clear testbed for measuring current LLM limitations and guiding future work on scientific reasoning in materials science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。