用生成数据缓解材料科学数据稀缺问题,效果接近甚至超过真实数据。
MatWheel: Addressing Data Scarcity in Materials Science Through Synthetic Data
- 用条件生成模型合成材料数据,支持全监督和半监督训练。
- 在两个数据稀缺数据集上,合成数据性能接近或超越真实数据。
- 适合材料属性预测研究者,尤其适用于数据获取困难的场景。
材料科学长期面临数据稀缺与标注成本高的挑战。受计算机视觉领域启发,我们提出MatWheel框架,利用条件生成模型生成的合成数据训练材料属性预测模型。探索了全监督与半监督两种学习场景。采用CGCNN进行属性预测,使用Con-CDVAE作为条件生成模型,在Matminer数据库中的两个数据稀缺材料属性数据集上进行了实验。结果表明,合成数据在极端数据稀缺情况下具有潜力,两项任务的表现均接近或超过真实样本。此外,伪标签对生成数据质量影响较小。未来工作将引入更先进模型并优化生成条件,以提升材料数据飞轮的效果。
原文摘要 · Abstract (English)
Data scarcity and the high cost of annotation have long been persistent challenges in the field of materials science. Inspired by its potential in other fields like computer vision, we propose the MatWheel framework, which train the material property prediction model using the synthetic data generated by the conditional generative model. We explore two scenarios: fully-supervised and semi-supervised learning. Using CGCNN for property prediction and Con-CDVAE as the conditional generative model, experiments on two data-scarce material property datasets from Matminer database are conducted. Results show that synthetic data has potential in extreme data-scarce scenarios, achieving performance close to or exceeding that of real samples in all two tasks. We also find that pseudo-labels have little impact on generated data quality. Future work will integrate advanced models and optimize generation conditions to boost the effectiveness of the materials data flywheel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。