首个同时包含模式与事实的完整知识图谱数据集构建工具
Return of the Schema: Building Complete Datasets for Machine Learning and Reasoning on Knowledge Graphs
- 提出全流程方法,从源图谱中提取模式与事实
- 构建多个含丰富模式的新数据集,支持推理与机器学习
- 输出OWL格式并兼容主流机器学习框架
现有知识图谱评估数据集通常仅包含实体关系事实,缺乏完整的模式信息,限制了依赖本体约束、推理或神经符号技术的方法在真实场景下的评估。本文提出 esource{},首个可提取包含模式与事实的完整数据集的工作流,支持机器学习与推理服务。该流程能处理模式与事实间的不一致性,并利用推理补全隐含知识。所构建的数据集套件涵盖具有丰富模式的知识图谱,同时为已有数据集补充模式信息。每个数据集以OWL格式存储,可直接用于推理;并提供工具将数据转换为标准机器学习库所需的张量表示。
原文摘要 · Abstract (English)
Datasets for the experimental evaluation of knowledge graph refinement algorithms typically contain only ground facts, retaining very limited schema level knowledge even when such information is available in the source knowledge graphs. This limits the evaluation of methods that rely on rich ontological constraints, reasoning or neurosymbolic techniques and ultimately prevents assessing their performance in large-scale, real-world knowledge graphs. In this paper, we present \resource{} the first resource that provides a workflow for extracting datasets including both schema and ground facts, ready for machine learning and reasoning services, along with the resulting curated suite of datasets. The workflow also handles inconsistencies detected when keeping both schema and facts and also leverage reasoning for entailing implicit knowledge. The suite includes newly extracted datasets from KGs with expressive schemas while simultaneously enriching existing datasets with schema information. Each dataset is serialized in OWL making it ready for reasoning services. Moreover, we provide utilities for loading datasets in tensor representations typical of standard machine learning libraries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。