arXiv:2602.17730physics.chem-phcond-mat.mtrl-sci2026-02被引 1

模型预测材料性能可能靠文献信息而非化学原理,需警惕假象。

Clever Materials: When Models Identify Good Materials for the Wrong Reasons

  • 用文献元数据替代化学特征,模型仍能准确预测材料性能
  • 五个材料任务中,模型对发表信息的预测远超随机水平
  • 研究呼吁引入检验机制,区分预测能力与真实化学理解

机器学习可加速材料发现,现有模型在多个基准上表现优异。然而,优秀表现未必意味着模型真正学习了化学规律。本文提出一个具体替代假设:属性预测可能受文献关联性干扰驱动。在五类任务中——金属有机框架(热/溶剂稳定性)、钙钛矿太阳能电池(效率)、电池(容量)和TADF发光体(发射波长)——使用标准化学描述符训练的模型,对作者、期刊和发表年份的预测均显著优于随机水平。当仅以这些被预测出的文献元数据(‘文献指纹’)作为输入时,第二个模型的表现有时甚至可媲美基于传统描述符的预测器。这表明,许多数据集无法排除非化学因素导致成功的可能性。研究进展需要常态化进行反证测试(如按组或时间划分数据、删除元数据),设计抗虚假相关性的数据集,并明确区分预测效用与化学理解的证据。

原文摘要 · Abstract (English)

Machine learning can accelerate materials discovery. Models perform impressively on many benchmarks. However, strong benchmark performance does not imply that a model learned chemistry. I test a concrete alternative hypothesis: that property prediction can be driven by bibliographic confounding. Across five tasks spanning MOFs (thermal and solvent stability), perovskite solar cells (efficiency), batteries (capacity), and TADF emitters (emission wavelength), models trained on standard chemical descriptors predict author, journal, and publication year well above chance. When these predicted metadata ("bibliographic fingerprints") are used as the sole input to a second model, performance is sometimes competitive with conventional descriptor-based predictors. These results show that many datasets do not rule out non-chemical explanations of success. Progress requires routine falsification tests (e.g., group/time splits and metadata ablations), datasets designed to resist spurious correlations, and explicit separation of two goals: predictive utility versus evidence of chemical understanding.

材料发现机器学习元数据偏差模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。