arXiv:2505.03049cs.LGcond-mat.mtrl-sci2025-05被引 9

34个案例展示大模型如何加速材料与化学研究全流程

34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery

  • 通过34个项目展示大模型在分子设计、数据管理等7大领域的应用
  • 在低数据环境下表现优异,支持跨学科研究与快速原型开发
  • 适合科研人员、自动化工具开发者及科学智能探索者

大型语言模型(LLMs)正在重塑材料科学与化学研究的多个方面,推动分子性质预测、材料设计、科研自动化、知识提取等进展。本文回顾了第二届全球混合式材料与化学大模型黑客松活动中开发的34个项目,涵盖七大关键领域:(1) 分子与材料性质预测,(2) 分子与材料设计,(3) 自动化与新型交互界面,(4) 科学传播与教育,(5) 研究数据管理与自动化,(6) 假设生成与评估,(7) 科学文献中的知识提取与推理。这些应用表明,大模型可作为多功能预测工具、领域专用工具的快速原型平台。特别是,通过引入推理能力、更多训练数据及新方法,开源与专有大模型性能提升,在低数据环境和跨学科研究中尤为有效。随着模型持续进步,其融入科研流程既带来新机遇也引发可靠性、可解释性与可复现性挑战,亟需持续探索与深入研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 34 total projects developed during the second annual Large Language Model Hackathon for Applications in Materials Science and Chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

大模型材料科学科研自动化智能助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。