arXiv:2502.16457cs.CL2025-02被引 5

构建1.7万条合成数据集,用大模型自动评估材料合成方案。

Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge

  • 基于1.7万条专家验证的合成配方构建新基准
  • 大模型评价与专家打分高度一致,支持自动化评估
  • 适合材料合成、AI辅助实验设计的研究者

材料合成对能源存储、催化、电子及生物医学设备等创新至关重要,但目前仍依赖经验性试错和专家直觉。本文构建了一个包含17,000条来自开放文献的专家验证合成配方的数据集,形成新的基准平台AlchemyBench。该平台提供端到端框架,支持大语言模型在原料预测、设备选择、合成流程生成及表征结果预判等任务中的研究。我们提出一种LLM-as-a-Judge评估框架,利用大语言模型实现自动化评价,其统计结果与专家评估高度一致。本工作为探索大模型在材料合成预测与指导中的能力提供了坚实基础,有助于提升实验设计效率,加速材料科学创新。

原文摘要 · Abstract (English)

Materials synthesis is vital for innovations such as energy storage, catalysis, electronics, and biomedical devices. Yet, the process relies heavily on empirical, trial-and-error methods guided by expert intuition. Our work aims to support the materials science community by providing a practical, data-driven resource. We have curated a comprehensive dataset of 17K expert-verified synthesis recipes from open-access literature, which forms the basis of our newly developed benchmark, AlchemyBench. AlchemyBench offers an end-to-end framework that supports research in large language models applied to synthesis prediction. It encompasses key tasks, including raw materials and equipment prediction, synthesis procedure generation, and characterization outcome forecasting. We propose an LLM-as-a-Judge framework that leverages large language models for automated evaluation, demonstrating strong statistical agreement with expert assessments. Overall, our contributions offer a supportive foundation for exploring the capabilities of LLMs in predicting and guiding materials synthesis, ultimately paving the way for more efficient experimental design and accelerated innovation in materials science.

材料发现大模型评估合成预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。