arXiv:2511.15222cond-mat.mtrl-scics.LG2025-11被引 3

用声子信息生成数据,让机器学习更准预测材料性能。

Why Physics Still Matters: Improving Machine Learning Prediction of Material Properties with Phonon-Informed Datasets

  • 用晶格振动指导采样构建训练集,提升数据代表性。
  • 声子引导模型在更少数据下仍优于随机数据训练的模型。
  • 适合材料性能预测、低对称性结构研究的科研人员参考。

机器学习已能以接近第一性原理的精度高效预测材料性质,但其性能高度依赖训练数据的质量、规模与多样性。在材料科学中,学习低对称性原子构型(如热激发、结构缺陷和化学无序)尤为关键,而这些特征在多数数据集中未充分覆盖。本文对比了图神经网络(GNN)在两类数据集上的表现:一类为随机生成的原子构型,另一类基于声子信息进行物理引导采样。以典型光电子材料在真实有限温度下的电子与力学性质预测为例,结果表明,尽管样本量更少,声子引导模型始终优于随机数据训练的模型。可解释性分析显示,高性能模型更关注控制性质变化的化学键,凸显物理引导数据生成的重要性。本工作证明,更大数据集未必带来更好模型,提出一种简单通用的数据构建策略,适用于材料信息学中的高质量训练数据生成。

原文摘要 · Abstract (English)

Machine learning (ML) methods have become powerful tools for predicting material properties with near first-principles accuracy and vastly reduced computational cost. However, the performance of ML models critically depends on the quality, size, and diversity of the training dataset. In materials science, this dependence is particularly important for learning from low-symmetry atomistic configurations that capture thermal excitations, structural defects, and chemical disorder, features that are ubiquitous in real materials but underrepresented in most datasets. The absence of systematic strategies for generating representative training data may therefore limit the predictive power of ML models in technologically critical fields such as energy conversion and photonics. In this work, we assess the effectiveness of graph neural network (GNN) models trained on two fundamentally different types of datasets: one composed of randomly generated atomic configurations and another constructed using physically informed sampling based on lattice vibrations. As a case study, we address the challenging task of predicting electronic and mechanical properties of a prototypical family of optoelectronic materials under realistic finite-temperature conditions. We find that the phonons-informed model consistently outperforms the randomly trained counterpart, despite relying on fewer data points. Explainability analyses further reveal that high-performing models assign greater weight to chemically meaningful bonds that control property variations, underscoring the importance of physically guided data generation. Overall, this work demonstrates that larger datasets do not necessarily yield better GNN predictive models and introduces a simple and general strategy for efficiently constructing high-quality training data in materials informatics.

材料预测图神经网络声子信息数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。