联邦学习让化工企业合作训练模型,不共享数据也能提升预测精度。
Federated Learning from Molecules to Processes: A Perspective
- 多方协作训练模型,数据不出本地,保护企业隐私。
- 在混合物活度系数和精馏塔建模中,联邦模型精度显著优于单方独立训练。
- 适合需要数据保密的化工企业联合开发智能模型的场景。
我们提出了化学工程领域联邦学习的展望,旨在推动化工行业机器学习的协同研发。大量化学与过程数据被化工企业私有化,形成数据孤岛,阻碍了大规模数据集上的机器学习模型训练。近年来,联邦学习因其可在不共享原始数据的前提下实现多方协同建模而受到关注。本文探讨了联邦学习在分子至过程尺度多个化工领域的应用潜力,并通过两个典型案例模拟多企业持有多源私有数据的实际场景:(i)基于图神经网络预测二元混合物活度系数;(ii)利用自编码器进行精馏塔系统辨识。结果表明,联邦学习联合训练的模型精度显著高于各企业独立训练的模型,且接近于合并所有企业数据后训练的集中式模型。因此,联邦学习在保障企业数据隐私的同时,极大提升了化工领域机器学习模型的性能,具有广阔工业应用前景。
原文摘要 · Abstract (English)
We present a perspective on federated learning in chemical engineering that envisions collaborative efforts in machine learning (ML) developments within the chemical industry. Large amounts of chemical and process data are proprietary to chemical companies and are therefore locked in data silos, hindering the training of ML models on large data sets in chemical engineering. Recently, the concept of federated learning has gained increasing attention in ML research, enabling organizations to jointly train machine learning models without disclosure of their individual data. We discuss potential applications of federated learning in several fields of chemical engineering, from the molecular to the process scale. In addition, we apply federated learning in two exemplary case studies that simulate practical scenarios of multiple chemical companies holding proprietary data sets: (i) prediction of binary mixture activity coefficients with graph neural networks and (ii) system identification of a distillation column with autoencoders. Our results indicate that ML models jointly trained with federated learning yield significantly higher accuracy than models trained by each chemical company individually and can perform similarly to models trained on combined datasets from all companies. Federated learning has therefore great potential to advance ML models in chemical engineering while respecting corporate data privacy, making it promising for future industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。