arXiv:2605.22304cs.AIcs.DB2026-05被引 2

为知识图谱数据集成管道提供可量化评估基准

Evaluation of Pipelines for Data Integration into Knowledge Graphs

论文配图:Evaluation of Pipelines for Data Integration into Knowledge Graphs
图 1 · 摘自论文原文
  • 构建KGI-Bench基准,统一评估数据集成管道性能
  • 用覆盖率、正确性、一致性三指标衡量更新后图谱质量
  • 提供电影领域标准数据集,适合研究者对比不同集成方案

将新数据整合到知识图谱(KG)通常涉及多个任务组成的流水线。针对特定集成问题存在多种可能的流水线,但目前尚无通用方法来评估这些流水线的整体质量和性能,难以确定最优选择。为此,我们提出一个新的基准 KGI-Bench,用于评估将不同类型输入数据集成到现有知识图谱中的流水线。通过分析输出结果——即更新后的知识图谱——采用三个互补的质量度量:覆盖率、正确性和一致性。我们还提供了电影领域的基准数据集(种子知识图谱、三种格式的重叠输入数据、参考知识图谱作为真实标签)。为验证该基准的适用性和实用性,我们对12种流水线进行了比较评估,并分析了它们在不同输入数据格式和设计选择下的表现。

原文摘要 · Abstract (English)

Integrating new data into knowledge graphs (KG) typically involves different tasks that are executed within workflows or pipelines There are many possible pipelines for a specific integration problem but there is not yet a general approach to evaluate the overall quality and performance of such pipelines to be able to determine the best choices. We therefore propose a new benchmark KGI-Bench to evaluate integration pipelines that ingest different kinds of input data into an existing KG. We evaluate pipelines by analyzing their output, i.e., the updated KG, with the three complementary quality metrics coverage, correctness and consistency. We also provide benchmark datasets (seed KG, overlapping input data of three formats, reference KG as a ground truth) for the movie domain. To demonstrate the applicability and usefulness of the proposed benchmark, we comparatively evaluate 12 pipelines and analyze their behavior across different input data formats and design choices.

知识图谱数据集成评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。