材料生成模型的评估需更严谨,否则结果误导。
All that structure matches does not glitter
- 清理数据集重复结构,避免错误评价
- 发现碳-24数据集仅40%为唯一结构
- 提出新指标与分组方式,提升评估可信度
针对无机晶体生成模型的理论预测潜力,本文批判性审视了常见数据集与评估指标。研究发现:第一,材料数据集应包含唯一晶体结构,如碳-24数据集实际仅有约40%独特结构;第二,若多种化学组成的多型体数量众多,则不应随机划分数据集,如钙钛矿-5与MP-20数据集即存在此问题;第三,若不考虑相同构建块的结构多样性,仅报告匹配率会误导评估。为此,本文提出多项改进:提供去重后的碳-24版本,按原子数$N$分组、包含对映异构体的版本,以及仅单元格不同但结构相同的两个版本;为含多型体数据集设计新划分方式,确保多型体在同一子集内;并引入METRe与cRMSE两项新指标,以修正传统匹配率的缺陷。
原文摘要 · Abstract (English)
Generative models for materials, especially inorganic crystals, hold potential to transform the theoretical prediction of novel compounds and structures. Advancement in this field depends on robust benchmarks and minimal, information-rich datasets that enable meaningful model evaluation. This paper critically examines common datasets and reported metrics for a crystal structure prediction task$\unicode{x2014}$generating the most likely structures given the chemical composition of a material. We focus on three key issues: First, materials datasets should contain unique crystal structures; for example, we show that the widely-utilized carbon-24 dataset only contains $\approx$40% unique structures. Second, materials datasets should not be split randomly if polymorphs of many different compositions are numerous, which we find to be the case for the perov-5 and MP-20 datasets. Third, benchmarks can mislead if used uncritically, e.g., reporting a match rate metric without considering the structural variety exhibited by identical building blocks. To address these oft-overlooked issues, we introduce several fixes. We provide revised versions of the carbon-24 dataset: one with duplicates removed, one deduplicated and split by number of atoms $N$, one with enantiomorphs, and two containing only identical structures but with different unit cells. We also propose new splits for datasets with polymorphs, ensuring that polymorphs are grouped within each split subset, setting a more sensible standard for benchmarking model performance. Finally, we present METRe and cRMSE, new model evaluation metrics that can correct existing issues with the match rate metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。