用真实实验数据评估机器学习力场,发现多数模型在现实场景下表现不佳。
UniFFBench: Evaluating Universal Machine Learning Force Fields Against Experimental Measurements
- 构建包含1500+矿物的MinX数据集,覆盖85种元素和极端条件
- 6个顶尖力场在实验数据上平均密度误差超实用门槛
- 模型性能与训练数据覆盖度相关,而非建模方法本身
通用机器学习力场(UMLFFs)有望通过快速原子模拟推动材料科学进步,但现有评估仅依赖计算基准,难以反映真实性能。本文提出UniFFBench评估框架,包含MinX数据集——一个涵盖1500多个矿物系统、85种元素、极端热力学条件(0–5000 K,0–1000 GPa)及结构复杂性(如部分占据和无序)的多样化数据集。该数据集结合实验参考值,可有效评估UMLFF在化学空间和条件上的泛化能力,远超典型训练范围。对六种先进UMLFF的系统评估揭示显著的“现实差距”:在计算基准上表现优异的模型,在面对实验复杂性时往往失效。即使最佳模型的密度预测误差也超过实际应用所需阈值。我们还发现模拟稳定性与力学性质准确性之间存在脱节,预测误差与训练数据代表性相关,而非建模方法本身。
原文摘要 · Abstract (English)
Universal machine learning force fields (UMLFFs) promise to revolutionize materials science by enabling rapid atomistic simulations across the periodic table. However, their evaluation has been limited to computational benchmarks that may not reflect real-world performance. We introduce UniFFBench, a comprehensive evaluation framework featuring the MinX dataset -- a diverse collection of 1,500+ mineral systems spanning 85 elements, extreme thermodynamic conditions (0--5000 K, 0--1000 GPa), and structural complexity, including partial occupancy and disorder. This diversity, combined with experimental reference values for validation, enables assessment of UMLFF generalization across chemical space and conditions substantially beyond typical training scenarios. Our systematic evaluation of six state-of-the-art UMLFFs reveals a substantial ``reality gap'': models achieving impressive performance on computational benchmarks often fail when confronted with experimental complexity. Even the best-performing models exhibit higher density prediction error than the threshold required for practical applications. We observe disconnects between simulation stability and mechanical property accuracy, with prediction errors correlating with training data representation rather than the modeling method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。