arXiv:2507.17980physics.comp-phcs.LG2025-07被引 3

用机器学习从分子模拟中自动提取聚合物结晶特征,精准量化结晶度。

Machine Learning Workflow for Analysis of High-Dimensional Order Parameter Space: A Case Study of Polymer Crystallization from Molecular Dynamics Simulations

  • 构建高维原子特征向量,结合几何、热力学与对称性描述符
  • 仅用3个序参量即可实现超过0.98的结晶分类准确率(AUC)
  • 提出可实时计算的结晶指数,适合大规模模拟中的结构演化监测

当前聚合物结晶路径分析依赖于单个序参量在预设阈值下的判断,易受阈值敏感性和系统偏差影响。本研究提出一种集成机器学习工作流,基于原子尺度分子动力学数据精确量化聚合物体系的结晶度。每个原子用包含几何、类热力学和对称性描述符的高维特征向量表示,通过低维嵌入揭示原子环境的潜在结构指纹。无监督聚类在嵌入空间中高精度识别晶态与非晶态原子。基于高质量标签,采用监督学习筛选出能完全捕捉结晶标签的最小序参量集。实验表明,仅需三个序参量即可重建结晶标签。据此定义结晶指数(C-index)为逻辑回归模型的结晶概率,该指数在整个过程中保持双峰分布,分类性能超过0.98(AUC)。值得注意的是,仅用一两个快照训练的模型即可实现高效在线计算结晶度。最后,我们展示最优C-index在结晶各阶段的演变,支持早期成核以熵为主导、后期对称性起关键作用的假设。该工作流提供数据驱动的序参量选择策略与大尺度模拟中结构转变的监测指标。

原文摘要 · Abstract (English)

Currently, identification of crystallization pathways in polymers is being carried out using molecular simulation-based data on a preset cut-off point on a single order parameter (OP) to define nucleated or crystallized regions. Aside from sensitivity to cut-off, each of these OPs introduces its own systematic biases. In this study, an integrated machine learning workflow is presented to accurately quantify crystallinity in polymeric systems using atomistic molecular dynamics data. Each atom is represented by a high-dimensional feature vector that combines geometric, thermodynamic-like, and symmetry-based descriptors. Low dimensional embeddings are employed to expose latent structural fingerprints within atomic environments. Subsequently, unsupervised clustering on the embeddings identified crystalline and amorphous atoms with high fidelity. After generating high quality labels with multidimensional data, we use supervised learning techniques to identify a minimal set of order parameters that can fully capture this label. Various tests were conducted to reduce the feature set, demonstrating that using only three order parameters is sufficient to recreate the crystallization labels. Based on these observed OPs, the crystallinity index (C-index) is defined as the logistic regression model's probability of crystallinity, remaining bimodal throughout the process and achieving over 0.98 classification performance (AUC). Notably, a model trained on one or a few snapshots enables efficient on-the-fly computation of crystallinity. Lastly, we demonstrate how the optimal C-index fit evolves during various stages of crystallization, supporting the hypothesis that entropy dominates early nucleation, while symmetry gains relevance later. This workflow provides a data-driven strategy for OP selection and a metric to monitor structural transformations in large-scale polymer simulations.

聚合物结晶机器学习序参量分子模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。