用统计集合模拟聚合物,让机器学习更贴近真实材料特性
PolySet: Restoring the Statistical Ensemble Nature of Polymers for Machine Learning
- 将聚合物视为由分子量分布采样的加权链集合,而非单一分子图
- 能保留高阶分布矩(如Mz、Mz+1),提升对尾部敏感性质的预测精度
- 适用于各类聚合物结构,为未来复杂体系建模提供物理基础
聚合物科学中的机器学习模型通常将聚合物视为单一、精确的分子图,而真实材料是由具有分布长度的随机链组成的统计集合。这种物理现实与数字表示之间的偏差限制了现有模型对聚合物行为的捕捉能力。本文提出PolySet框架,将聚合物表示为从假设的摩尔质量分布中采样的有限加权链集合。该集合编码方式不依赖化学细节,兼容任意分子表示,本文以均聚物为例,结合最小语言模型进行演示。结果表明,PolySet能够保留高阶分布矩(如Mz、Mz+1),使机器学习模型在预测尾部敏感性质时具备显著更高的稳定性和准确性。通过显式承认聚合物物质的统计本质,PolySet为未来的聚合物机器学习建立了一个物理上合理的基础,且天然可扩展至共聚物、嵌段结构及其他复杂拓扑。
原文摘要 · Abstract (English)
Machine-learning (ML) models in polymer science typically treat a polymer as a single, perfectly defined molecular graph, even though real materials consist of stochastic ensembles of chains with distributed lengths. This mismatch between physical reality and digital representation limits the ability of current models to capture polymer behaviour. Here we introduce PolySet, a framework that represents a polymer as a finite, weighted ensemble of chains sampled from an assumed molar-mass distribution. This ensemble-based encoding is independent of chemical detail, compatible with any molecular representation and illustrated here in the homopolymer case using a minimal language model. We show that PolySet retains higher-order distributional moments (such as Mz, Mz+1), enabling ML models to learn tail-sensitive properties with greatly improved stability and accuracy. By explicitly acknowledging the statistical nature of polymer matter, PolySet establishes a physically grounded foundation for future polymer machine learning, naturally extensible to copolymers, block architectures, and other complex topologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。