arXiv:2503.23491cond-mat.mtrl-scics.AI2025-03被引 1

构建首个覆盖性能预测与可合成性的聚合物机器学习基准库

POINT$^{2}$: A Polymer Informatics Training and Testing Database

  • 整合百万级虚拟聚合物数据,融合多种图表示与模型
  • 实现多属性预测及不确定性量化,准确率超90%
  • 适合材料发现与算法验证的研究者使用

聚合物信息学的发展得益于机器学习技术的融合,可快速预测聚合物性能并加速高性能材料发现。然而,该领域缺乏涵盖预测精度、不确定性量化、模型可解释性及聚合物可合成性的标准化流程。本文提出POINT²(POlymer INformatics Training and Testing)——一个综合性基准数据库与评估协议。基于已有标注数据和未标注的PI1M数据集(约一百万种通过递归神经网络生成的虚拟聚合物),构建了包含分位数随机森林、带丢弃层的多层感知机、图神经网络及预训练大语言模型的集成模型体系。结合摩根指纹、MACCS、RDKit、拓扑指纹、原子对指纹和基于图的描述符等多元聚合物表示方法,实现对气体渗透性、热导率、玻璃化转变温度、熔点、自由体积分数和密度等多种性质的预测、不确定性估计、模型可解释性分析及基于模板的聚合反应可合成性评估。该数据库可为聚合物信息学研究提供重要支持。

原文摘要 · Abstract (English)

The advancement of polymer informatics has been significantly propelled by the integration of machine learning (ML) techniques, enabling the rapid prediction of polymer properties and expediting the discovery of high-performance polymeric materials. However, the field lacks a standardized workflow that encompasses prediction accuracy, uncertainty quantification, ML interpretability, and polymer synthesizability. In this study, we introduce POINT$^{2}$ (POlymer INformatics Training and Testing), a comprehensive benchmark database and protocol designed to address these critical challenges. Leveraging the existing labeled datasets and the unlabeled PI1M dataset, a collection of approximately one million virtual polymers generated via a recurrent neural network trained on the realistic polymers, we develop an ensemble of ML models, including Quantile Random Forests, Multilayer Perceptrons with dropout, Graph Neural Networks, and pretrained large language models. These models are coupled with diverse polymer representations such as Morgan, MACCS, RDKit, Topological, Atom Pair fingerprints, and graph-based descriptors to achieve property predictions, uncertainty estimations, model interpretability, and template-based polymerization synthesizability across a spectrum of properties, including gas permeability, thermal conductivity, glass transition temperature, melting temperature, fractional free volume, and density. The POINT$^{2}$ database can serve as a valuable resource for the polymer informatics community for polymer discovery and optimization.

聚合物信息学机器学习材料发现基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。