arXiv:2412.04780cs.LG2024-12被引 1

提出无监督方法SEKA检测知识图谱中的异常三元组与实体,提升数据质量。

Anomaly Detection and Classification in Knowledge Graphs

  • 基于路径排名算法改进的CPRA,高效识别知识图谱异常
  • 在四个真实数据集上优于基线方法,有效发现冗余、矛盾等异常
  • 构建异常类型分类体系TAXO,帮助理解知识图谱数据质量问题

知识图谱中不可避免地存在冗余、不一致、矛盾和缺失等异常,因其常通过人工或机器学习技术构建。本文提出SEKA(SEeking Knowledge graph Anomalies),一种无监督的异常三元组与实体检测方法,可在保持知识图谱覆盖范围的同时提升其准确性。我们改进了路径排名算法(PRA),提出协同路径排名算法(CPRA),专门用于检测知识图谱中的异常。同时,我们构建了TAXO(TAXOnomy of anomaly types in KGs)——一种知识图谱异常类型的分类体系,对SEKA发现的异常进行系统分类,并深入讨论知识图谱中可能存在的数据质量问题。我们在四个真实世界知识图谱(YAGO-1、KBpedia、Wikidata、DSKG)上评估了SEKA和TAXO,结果表明其性能显著优于基线方法。

原文摘要 · Abstract (English)

Anomalies such as redundant, inconsistent, contradictory, and deficient values in a Knowledge Graph (KG) are unavoidable, as these graphs are often curated manually, or extracted using machine learning and natural language processing techniques. Therefore, anomaly detection is a task that can enhance the quality of KGs. In this paper, we propose SEKA (SEeking Knowledge graph Anomalies), an unsupervised approach for the detection of abnormal triples and entities in KGs. SEKA can help improve the correctness of a KG whilst retaining its coverage. We propose an adaption of the Path Rank Algorithm (PRA), named the Corroborative Path Rank Algorithm (CPRA), which is an efficient adaptation of PRA that is customized to detect anomalies in KGs. Furthermore, we also present TAXO (TAXOnomy of anomaly types in KGs), a taxonomy of possible anomaly types that can occur in a KG. This taxonomy provides a classification of the anomalies discovered by SEKA with an extensive discussion of possible data quality issues in a KG. We evaluate both approaches using the four real-world KGs YAGO-1, KBpedia, Wikidata, and DSKG to demonstrate the ability of SEKA and TAXO to outperform the baselines.

知识图谱异常检测数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。