用大模型+本体技术自动提取材料计算流程,让论文结果可复现可比较。
Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction
- 用提示工程+多阶段过滤,从文献中提取密度泛函计算流程和参数。
- 构建了对齐材料本体的知識圖譜,支持堆垛层错能数据系统比對。
- 适合材料计算、数据复现与知识图谱研究者使用。
材料科学中的计算结果复现仍面临挑战,因模拟流程与参数常以非结构化文本或表格形式呈现。尽管文献数据对验证和复用有价值,但缺乏机器可读的流程描述,限制了大规模整理与系统性比较。现有文本挖掘方法难以完整提取包含参数的计算流程。本文提出一种基于本体驱动、大语言模型辅助的框架,实现从文献中自动化提取计算流程。聚焦于密排六方镁及其二元合金的密度泛函理论堆垛层错能(SFE)计算,采用多阶段过滤策略结合提示工程的LLM抽取方法,应用于方法部分与表格。提取信息统一为标准模式,并对齐材料本体(CMSO、ASMO、PLDO),通过atomRDF构建知识图谱。该图谱支持已报道SFE值的系统比较,推动计算协议的结构化复用。尽管完全复现仍受限于缺失或隐含元数据,该框架为组织与语义互操作化发表结果提供了基础,提升了计算材料数据的透明度与可复用性。
原文摘要 · Abstract (English)
Reproducibility of computational results remains a challenge in materials science, as simulation workflows and parameters are often reported only in unstructured text and tables. While literature data are valuable for validation and reuse, the lack of machine-readable workflow descriptions prevents large-scale curation and systematic comparison. Existing text-mining approaches are insufficient to extract complete computational workflows with their associated parameters. An ontology-driven, large language model (LLM)-assisted framework is introduced for the automated extraction and structuring of computational workflows from the literature. The approach focuses on density functional theory-based stacking fault energy (SFE) calculations in hexagonal close-packed magnesium and its binary alloys, and uses a multi-stage filtering strategy together with prompt-engineered LLM extraction applied to method sections and tables. Extracted information is unified into a canonical schema and aligned with established materials ontologies (CMSO, ASMO, and PLDO), enabling the construction of a knowledge graph using atomRDF. The resulting knowledge graph enables systematic comparison of reported SFE values and supports the structured reuse of computational protocols. While full computational reproducibility is still constrained by missing or implicit metadata, the framework provides a foundation for organizing and contextualizing published results in a semantically interoperable form, thereby improving transparency and reusability of computational materials data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。