用机器学习预测自由基共聚单体反应性比率,兼顾化学洞察与精准预测。
Chemically-Informed Machine Learning Approach for Prediction of Reactivity Ratios in Radical Copolymerization
- 通过谱聚类将单体按物理化学特征分三组,发现不同组间反应模式差异。
- 神经网络模型在完整数据集上表现更优,但特定领域训练提升局部精度。
- 适合材料设计初期探索与后期精准预测,根据数据量灵活选择策略。
预测单体反应性比率对于调控共聚物的单体序列分布及其性能至关重要。传统实验方法耗时且资源密集,现有计算方法常面临准确率或可扩展性不足的问题。本文提出一种结合无监督学习与人工神经网络的方法,用于预测自由基共聚中的反应性比率。通过对单体的物理化学特征进行谱聚类,识别出三类具有特征反应模式的单体群组。该计算高效聚类方法揭示了不同单体组间的相互作用,导致交替、随机、嵌段和梯度共聚物等不同序列结构,为初始探索提供化学洞察。基于这些发现,我们训练人工神经网络实现定量反应性比率预测。研究比较了直接特征拼接与分组特化训练两种融合策略,后者在特定化学领域表现出性能提升,但使用完整数据集的模型整体表现更优,揭示了化学特异性与数据可用性之间的根本权衡。本工作表明,无监督学习适用于快速化学洞察探索,而有监督学习则满足最终设计预测所需的精度,最优策略取决于数据可用性和应用需求。
原文摘要 · Abstract (English)
Predicting monomer reactivity ratios is crucial for controlling monomer sequence distribution in copolymers and their properties. Traditional experimental methods of determining reactivity ratios are time-consuming and resource-intensive, while existing computational methods often struggle with accuracy or scalability. Here, we present a method that combines unsupervised learning with artificial neural networks to predict reactivity ratios in radical copolymerization. By applying spectral clustering to physicochemical features of monomers, we identified three distinct monomer groups with characteristic reactivity patterns. This computationally efficient clustering approach revealed specific monomer group interactions leading to different sequence arrangements, including alternating, random, block, and gradient copolymers, providing chemical insights for initial exploration. Building upon these insights, we trained artificial neural networks to achieve quantitative reactivity ratio predictions. We explored two integration strategies including direct feature concatenation, and cluster-specific training, which demonstrated performance enhancements for targeted chemical domains compared to general training with equivalent sample sizes. However, models utilizing complete datasets outperformed specialized models trained on focused subsets, revealing a fundamental trade-off between chemical specificity and data availability. This work demonstrates that unsupervised learning offers rapid chemical insight for exploratory analysis, while supervised learning provides the accuracy necessary for final design predictions, with optimal strategies depending on data availability and application requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。