为机器学习势函数建立标准化元数据框架,助力可复现与比较。
An Ontology for Machine Learning Interatomic Potentials
- 构建了涵盖方法、数据、基准的三模块本体,统一描述MLIP研究要素
- 通过27条形式化公理确保数据完整一致,支持模型与训练数据的关联追溯
- 适用于材料模拟研究者、数据共享平台及机器学习势函数开发者
机器学习势函数(MLIPs)以远低于密度泛函理论(DFT)或波函数方法的成本近似量子力学能量与力。该领域包含日益增长的算法、训练数据集、超参数和目标材料,但系统比较、复现和延续研究所需的元数据仍分散在论文、脚本和非标准文件格式中。本文提出MLIPs本体(OWL 2 DL),涵盖描述MLIP方法、超参数、带有DFT溯源的训练数据集以及已发表基准测试所需概念。本体分为方法、训练数据和基准三个模块,整合材料科学(MDO、CMSO/ASMO)与机器学习(ML-Schema)现有本体,并补充如Croissant等数据集侧模式。其包含27条正式公理,强制数据完整性与一致性,包括属性链,将训练模型与其方法和训练数据关联。通过基于矩张量势的实例演示,结合20篇论文构建的知识图谱进行能力问题执行、OWL推理及与现有本体的对比评估,验证了其有效性。
原文摘要 · Abstract (English)
Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost. The field encompasses a growing ecosystem of algorithms, training datasets, hyperparameters, and target materials, yet the metadata needed to systematically compare, reproduce, and build upon MLIP studies remains scattered across papers, scripts, and ad-hoc file formats. We present the MLIPs ontology, an OWL 2 DL ontology that captures the concepts needed to describe MLIP methods, their hyperparameters, training datasets with DFT provenance, and published benchmarks. The ontology is organized into three modules---Method, Training Data, and Benchmark---and connects existing ontologies in materials science (MDO, CMSO/ASMO) and machine learning (ML-Schema), complementing dataset-side schemas such as Croissant. It declares 27 formal axioms enforcing data completeness and consistency, including property chains that link trained models to their methods and training data. We demonstrate the ontology through a running example based on Moment Tensor Potentials and evaluate it through competency-question execution on a 20-paper seeded knowledge graph, OWL reasoning, and comparison with existing ontologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。