arXiv:2507.07293cond-mat.mtrl-scics.LG2025-07被引 1

用大模型自动挖文献数据,预测矿物热力学参数。

Thermodynamic Prediction Enabled by Automatic Dataset Building and Machine Learning

  • 用大模型自动提取文献中的化学数据,构建机器可读数据库。
  • 基于自动生成的数据训练模型,准确预测矿物生成焓等参数。
  • 适合材料与化学领域研究者,加速新物质发现。

化学与材料科学的新发现日益增多,所需知识量和实验工作量持续上升,为机器学习(ML)加速科研效率提供了独特机遇。本文展示:(1)利用大语言模型(LLMs)实现自动化文献综述;(2)训练机器学习模型预测化学知识(热力学参数)。基于LLM的文献分析工具LMExt成功将金属阳离子-配体稳定常数、热力学性质等化学信息及其他类型数据(如医学论文、金融报告)提取为机器可读结构,有效克服各领域固有挑战。借助自主获取的热力学数据,采用CatBoost算法训练的模型实现了对矿物生成焓等热力学参数的精准预测。本工作凸显了集成式机器学习方法在重塑化学与材料科学研究方面的变革潜力。

原文摘要 · Abstract (English)

New discoveries in chemistry and materials science, with increasingly expanding volume of requisite knowledge and experimental workload, provide unique opportunities for machine learning (ML) to take critical roles in accelerating research efficiency. Here, we demonstrate (1) the use of large language models (LLMs) for automated literature reviews, and (2) the training of an ML model to predict chemical knowledge (thermodynamic parameters). Our LLM-based literature review tool (LMExt) successfully extracted chemical information and beyond into a machine-readable structure, including stability constants for metal cation-ligand interactions, thermodynamic properties, and other broader data types (medical research papers, and financial reports), effectively overcoming the challenges inherent in each domain. Using the autonomous acquisition of thermodynamic data, an ML model was trained using the CatBoost algorithm for accurately predicting thermodynamic parameters (e.g., enthalpy of formation) of minerals. This work highlights the transformative potential of integrated ML approaches to reshape chemistry and materials science research.

热力学预测大模型数据挖掘材料发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。