用化学命名直接生成材料嵌入,无需结构数据也能预测性能
ReadMOF: Structure-Free Semantic Embeddings from Systematic MOF Nomenclature for Machine Learning
- 用预训练语言模型解析金属有机框架的系统命名,生成语义嵌入
- 在性质预测和相似性检索上达到与结构依赖方法相当的精度
- 适合无结构数据的材料发现,支持大模型进行化学推理
系统性化学名称(如金属有机框架的IUPAC命名)以标准化文本形式包含丰富的结构与组成信息。本文提出ReadMOF,据我们所知是首个无需原子坐标或连接图的命名无关机器学习框架,利用这些名称建模结构-性能关系。通过预训练语言模型,ReadMOF将剑桥结构数据库(CSD)中的系统性MOF名称转换为向量嵌入,其表征能力接近传统基于结构的描述符。该嵌入可应用于材料信息学任务,包括性质预测、相似性检索与聚类,性能与依赖几何的方法相当。结合大语言模型后,仅凭文本输入即可实现化学上的可解释推理。结果表明,通过现代自然语言处理技术解析的结构化化学语言,可提供一种可扩展、可解释且不依赖几何的分子表示替代方案,为材料科学中的语言驱动发现开辟新路径。
原文摘要 · Abstract (English)
Systematic chemical names, such as IUPAC-style nomenclature for metal-organic frameworks (MOFs), contain rich structural and compositional information in a standardized textual format. Here we introduce ReadMOF, which is, to our knowledge, the first nomenclature-free machine learning framework that leverages these names to model structure-property relationships without requiring atomic coordinates or connectivity graphs. By employing pretrained language models, ReadMOF converts systematic MOF names from the Cambridge Structural Database (CSD) into vector embeddings that closely represent traditional structure-based descriptors. These embeddings enable applications in materials informatics, including property prediction, similarity retrieval, and clustering, with performance comparable to geometry-dependent methods. When combined with large language models, ReadMOF also establishes chemically meaningful reasoning ability with textual input only. Our results show that structured chemical language, interpreted through modern natural language processing techniques, can provide a scalable, interpretable, and geometry-independent alternative to conventional molecular representations. This approach opens new opportunities for language-driven discovery in materials science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。