机器学习预测准晶近似相时,训练数据微小变化会导致预测结果完全不同。
On the Robustness of Machine Learning Models in Predicting Thermodynamic Properties: a Case of Searching for New Quasicrystal Approximants
- 构建嵌套的准晶近似相数据集,测试不同训练样本的影响。
- 不同合理数据变动使预测新材料集合完全改变,模型稳定性差。
- 提出预训练与顺序训练策略,显著提升预测一致性。
尽管基于人工智能的无序晶体建模是新材料设计中广泛使用且成熟的方法,但其鲁棒性、可靠性与稳定性问题仍未解决,甚至未得到充分讨论。为揭示该问题,本文构建了一系列嵌套的准晶近似相数据集,并在这些数据上训练多种机器学习模型。对预测结果的定性与定量评估清晰表明,训练样本的合理变化可导致完全不同的潜在新材料预测集合。此外,我们验证了预训练的优势,并提出一种简单有效的顺序训练方法,以增强模型预测的稳定性。
原文摘要 · Abstract (English)
Despite an artificial intelligence-assisted modeling of disordered crystals is a widely used and well-tried method of new materials design, the issues of its robustness, reliability, and stability are still not resolved and even not discussed enough. To highlight it, in this work we composed a series of nested intermetallic approximants of quasicrystals datasets and trained various machine learning models on them correspondingly. Our qualitative and, what is more important, quantitative assessment of the difference in the predictions clearly shows that different reasonable changes in the training sample can lead to the completely different set of the predicted potentially new materials. We also showed the advantage of pre-training and proposed a simple yet effective trick of sequential training to increase stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。