开源大模型助力材料科学,提升科研透明度与可复现性
Large language models in materials science and the need for open-source approaches
- 用大模型从文献中提取合成条件等关键信息
- 开源模型性能可媲美闭源模型,且成本更低
- 适合希望自主可控的科研团队和开放协作平台
大语言模型正快速改变材料科学。本综述系统分析了其在材料发现全流程中的应用,聚焦三大方向:从科学文献中挖掘信息、构建预测模型、协调多智能体实验系统。研究表明,LLM可自动提取合成条件,学习结构-性能关系,并整合计算工具与实验室自动化设备实现智能协作。尽管当前进展主要依赖闭源商业模型,但我们的基准测试显示,开源模型在性能上已能比肩,同时具备更高的透明度、可复现性、成本效益和数据隐私保护优势。随着开源模型持续进化,我们倡导更广泛采用,以构建开放、灵活、社区驱动的科学发现AI平台。
原文摘要 · Abstract (English)
Large language models (LLMs) are rapidly transforming materials science. This review examines recent LLM applications across the materials discovery pipeline, focusing on three key areas: mining scientific literature , predictive modelling, and multi-agent experimental systems. We highlight how LLMs extract valuable information such as synthesis conditions from text, learn structure-property relationships, and can coordinate agentic systems integrating computational tools and laboratory automation. While progress has been largely dependent on closed-source commercial models, our benchmark results demonstrate that open-source alternatives can match performance while offering greater transparency, reproducibility, cost-effectiveness, and data privacy. As open-source models continue to improve, we advocate their broader adoption to build accessible, flexible, and community-driven AI platforms for scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。