用化学先验知识将计算数据高效迁移到实验,少样本也能高精度预测催化剂性能。
Transfer learning from first-principles calculations to experiments with chemistry-informed domain transformation
- 基于化学先验构建计算与实验数据的映射空间,实现异质域融合。
- 仅用不到10组实验数据,精度就达到全量训练模型(超100组)的十分之一。
- 适合材料实验数据稀缺场景,可大幅减少实验室试错次数。
模拟到现实(Sim2Real)迁移学习是一种通过利用计算数据知识来高效解决真实世界任务的机器学习技术,在材料科学中被视为缓解实验数据稀缺问题的有前景方案。本文提出一种基于化学信息域变换的从第一性原理计算到实验的高效迁移学习方法,通过融合计算数据(源域)与实验数据(目标域)的异质信息,借助底层物理化学规律实现跨域整合。该方法将计算数据从模拟空间映射至实验空间,利用两类化学先验:(1) 统计系综特性,(2) 源域与目标域变量之间的关系。以逆水气变换反应催化剂活性预测为验证案例,结合大量第一性原理数据与少量实验数据进行演示。结果表明,迁移学习模型在准确性和数据效率方面均表现出正向迁移。尤其值得注意的是,仅使用少于10组目标数据进行域变换时,模型精度仍可达仅用超过100组目标数据训练的全量模型的约十分之一,说明该方法能以极低实验数据成本实现高性能预测,显著降低实际实验室中的试错次数。
原文摘要 · Abstract (English)
Simulation-to-Real (Sim2Real) transfer learning, the machine learning technique that efficiently solves a real-world task by leveraging knowledge from computational data, has received increasing attention in materials science as a promising solution to the scarcity of experimental data. We proposed an efficient transfer learning scheme from first-principles calculations to experiments based on the chemistry-informed domain transformation, that integrates the heterogeneous source and target domains by harnessing the underlying physics and chemistry. The proposed method maps the computational data from the simulation space (source domain) into the space of experimental data (target domain). During this process, these qualitatively different domains are efficiently integrated by a couple of prior knowledge of chemistry, (1) the statistical ensemble, and (2) the relationship between source and target quantities. As a proof-of-concept, we predict the catalyst activity for the reverse water-gas shift reaction by using the abundant first-principles data in addition to the experimental data. Through the demonstration, we confirmed that the transfer learning model exhibits positive transfer in accuracy and data efficiency. In particular, a significantly high accuracy was achieved despite using a few (less than ten) target data in domain transformation, whose accuracy is one order of magnitude smaller than that of a full scratch model trained with over 100 target data. This result indicates that the proposed method leverages the high prediction performance with few target data, which helps to save the number of trials in real laboratories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。