用少量标签+无标签晶体预测材料属性,提升精度与可靠性
Learning Materials Properties from Scarce Labels and Unlabeled Crystals

- 设计半监督基准SemiMat,统一评估框架与数据划分
- 提出MatRank方法,加权伪标签提升预测准确性(NMAE=0.896)
- 适合材料数据少、需可靠预测的研究者使用
从少量标注数据和大量无标签晶体中学习材料属性,是数据驱动材料发现的核心挑战。本文提出SemiMat,一个可控的半监督材料属性回归基准,涵盖六项稀疏标签任务、四种图骨干网络和五次预定义划分,固定输入、验证集选择、测试集报告及归一化平均绝对误差(NMAE)计算,并提供方法排名汇总。同时提出MatRank,通过标注锚点构建伪目标,依据局部可靠性与弱预测一致性加权,一致训练强弱图视图,并引入排序信号使无标签晶体同时影响预测值与候选顺序。在24个骨干-任务组合中,单一MatRank目标实现最低的保留测试集平均NMAE(0.896)和最佳平均方法排名(2.208)。通过分布外(OOD)、组件及生成池诊断,识别性能提升的可靠性区域与仍需进一步评估的环节。代码已开源。
原文摘要 · Abstract (English)
Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery. We present SemiMat, a controlled benchmark for semi-supervised materials property regression, and MatRank, a reliability-weighted objective for continuous pseudo-label uncertainty. SemiMat fixes labeled and unlabeled crystal inputs, graph-backbone interfaces, validation-only checkpoint selection, held-out test reporting, normalized MAE (NMAE), and method-rank summaries across six scarce-label tasks, four graph backbones, and five predefined split runs. MatRank builds pseudo-targets from labeled anchors, weights them by local reliability and weak-prediction agreement, trains weak and strong graph views consistently, and adds ranking signals so that unlabeled crystals shape both values and candidate order. Across the retained 24 backbone-task blocks, one fixed MatRank objective gives the lowest aggregate held-out test NMAE (0.896) and best average method rank (2.208). The component, OOD, and generated-pool diagnostics identify where the gain is reliable and where further screening evaluation remains necessary. Code is available at https://github.com/littlepeachs/SemiMat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。