用统一图结构表示化学过程,实现跨领域材料性质预测。
Accelerating Materials Discovery: Learning a Universal Representation of Chemical Processes for Cross-Domain Property Prediction
- 构建化学流程的有向树图表示,融合文本、分子结构与数据。
- 在近9000份文档的70万张流程图上训练,跨域迁移性能强。
- 小样本微调即可达成高精度,适合材料研发快速筛选。
化学过程的实验验证耗时且成本高昂,限制了材料发现的探索。机器学习可优先筛选有前景的候选物,但专利和文献中的数据异构且难以利用。本文提出一种通用的有向树状流程图表示法,将非结构化文本、分子结构和数值测量统一为单一可机读格式。为从该结构化数据中学习,我们开发了具备属性条件注意力机制的多模态图神经网络。模型在近9000份多样化文档中提取的约70万张流程图上训练,学习到语义丰富的嵌入表示,具备跨领域泛化能力。在少量特定领域数据上微调后,模型表现出优异性能,证明大规模预训练的通用流程表示可高效迁移至特定预测任务,仅需极少额外数据。
原文摘要 · Abstract (English)
Experimental validation of chemical processes is slow and costly, limiting exploration in materials discovery. Machine learning can prioritize promising candidates, but existing data in patents and literature is heterogeneous and difficult to use. We introduce a universal directed-tree process-graph representation that unifies unstructured text, molecular structures, and numeric measurements into a single machine-readable format. To learn from this structured data, we developed a multi-modal graph neural network with a property-conditioned attention mechanism. Trained on approximately 700,000 process graphs from nearly 9,000 diverse documents, our model learns semantically rich embeddings that generalize across domains. When fine-tuned on compact, domain-specific datasets, the pretrained model achieves strong performance, demonstrating that universal process representations learned at scale transfer effectively to specialized prediction tasks with minimal additional data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。