为实验数据建立统一语义标准,让不同实验室的数据能无缝互通。
The AnIML Ontology: Enabling Semantic Interoperability for Large-Scale Experimental Data in Interconnected Scientific Labs

- 用本体技术精确定义AnIML数据格式的语义,避免歧义。
- 通过知识图谱和SPARQL验证,确保数据解释一致可靠。
- 适合工业研发、跨实验室协作的研究者使用。
在异构实验数据系统间实现语义互操作性仍是科学发现数字化的重大障碍。分析信息标记语言(AnIML)作为基于XML的分析化学与生物学标准,正被工业研发实验室广泛用于管理与交换实验数据。然而,其灵活的XML模式允许不同利益相关方产生分歧解读,导致不一致,削弱了原本旨在支持的互操作性。本文提出AnIML本体(AnIML Ontology),一个基于OWL 2的本体,形式化定义AnIML语义,并与Allotrope数据格式对齐,以支持未来跨系统、跨实验室的互操作性。本体通过专家参与式方法开发,结合大模型辅助需求获取与协同本体工程。我们采用多层验证策略:将真实AnIML文件转换为知识图谱,通过SPARQL验证能力问题,并设计一种新型验证协议——基于对抗性负向能力问题映射到已知本体反模式,并通过SHACL约束强制执行。
原文摘要 · Abstract (English)
Achieving semantic interoperability across heterogeneous experimental data systems remains a major barrier to data-driven scientific discovery. The Analytical Information Markup Language (AnIML), a flexible XML-based standard for analytical chemistry and biology, is increasingly used in industrial R&D labs for managing and exchanging experimental data. However, the expressivity of the XML schema permits divergent interpretations across stakeholders, introducing inconsistencies that undermine the interoperability the AnIML schema was designed to support. In this paper, we present the AnIML Ontology, an OWL 2 ontology that formalises the semantics of AnIML and aligns it with the Allotrope Data Format to support future cross-system and cross-lab interoperability. The ontology was developed using an expert-in-the-loop approach combining LLM-assisted requirement elicitation with collaborative ontology engineering. We validate the ontology through a multi-layered approach: data-driven transformation of real-world AnIML files into knowledge graphs, competency question verification via SPARQL, and a novel validation protocol based on adversarial negative competency questions mapped to established ontological anti-patterns and enforced via SHACL constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。