arXiv:2409.09691cond-mat.softcond-mat.mtrl-sci2024-09

机器学习预测共聚物性质时,外推能力取决于数据量和范围。

Extrapolative ML Models for Copolymers

  • 用神经网络和XGBoost学习结构与性能的函数关系,提升外推能力。
  • 训练数据越多、覆盖范围越广,模型外推效果越好。
  • 适合材料设计中探索未知性能区域的研究者参考。

机器学习已广泛用于材料性质预测,可通过已有数据快速筛选庞大的物理化学空间。然而,现有模型本质上是内插式的,对已知性质范围之外的候选材料搜索能力尚不明确。模型性能与学习策略及训练数据量密切相关。本文研究了机器学习模型外推能力、训练数据规模与范围、以及学习方法之间的关系,聚焦于根据单体序列预测共聚物性质这一经典问题。发现基于树搜索的算法因依赖结构相似性,外推效率低下;而神经网络和XGBoost模型通过学习结构-性能间的函数关联,其外推能力与训练数据量和覆盖范围呈强相关。该结果对基于机器学习的新材料开发具有重要意义。

原文摘要 · Abstract (English)

Machine learning models have been progressively used for predicting materials properties. These models can be built using pre-existing data and are useful for rapidly screening the physicochemical space of a material, which is astronomically large. However, ML models are inherently interpolative, and their efficacy for searching candidates outside a material's known range of property is unresolved. Moreover, the performance of an ML model is intricately connected to its learning strategy and the volume of training data. Here, we determine the relationship between the extrapolation ability of an ML model, the size and range of its training dataset, and its learning approach. We focus on a canonical problem of predicting the properties of a copolymer as a function of the sequence of its monomers. Tree search algorithms, which learn the similarity between polymer structures, are found to be inefficient for extrapolation. Conversely, the extrapolation capability of neural networks and XGBoost models, which attempt to learn the underlying functional correlation between the structure and property of polymers, show strong correlations with the volume and range of training data. These findings have important implications on ML-based new material development.

机器学习材料预测外推能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。