构建物理化学推理基准,评估大模型材料合成假设生成能力
Mechanisms of Matter: Language Inferential Benchmark on Physicochemical Hypothesis in Materials Synthesis
- 设计八类纳米材料合成的物理化学推理基准
- 发现大模型推理脱离基本物化原理,准确率低
- 提出原则感知提示法,显著提升推理准确率与效率
大语言模型(LLMs)在生成材料合成科学假设方面的能力尚未得到量化评估,主要受限于缺乏探测物化逻辑推理的基准。为此,我们提出了MatterMech基准,用于评估LLM在八类纳米材料合成领域的假设生成能力。分析显示,尽管大模型擅长抽象逻辑,但其推理未能扎根于基本物理化学原理。我们证明,所提出的原理感知提示方法显著优于标准链式思维,大幅提升了假设准确率与计算效率。本工作为推进大模型在材料科学中可靠假设生成提供了方法论框架。MatterMech基准及代码已公开于GitHub。
原文摘要 · Abstract (English)
The capacity of Large Language Models (LLMs) to generate valid scientific hypotheses for materials synthesis remains largely unquantified, hindered by the absence of benchmarks probing physicochemical logics reasoning. To address this, we introduce MatterMech, a benchmark for evaluating LLM-generated hypotheses across eight nanomaterial synthesis domains. Our analysis reveals a critical disconnect: LLMs are proficient in abstract logic yet fail to ground their reasoning in fundamental physicochemical principles. We demonstrate that our proposed principle-aware prompting methodology substantially outperforms standard Chain-of-Thought, enhancing both hypothesis accuracy and computational efficiency. This work provides a methodological framework to advance LLMs toward reliable scientific hypothesis generation in materials science. The MatterMech benchmark and associated code is publicly available at \href{https://github.com/amair-lab/MatterMech}{GitHub}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。