arXiv:2511.18364cs.AIcs.DB2025-11被引 3

构建可复现的知识图谱集成管道,支持多种数据源与大模型协同

KGpipe: Generation and Evaluation of Pipelines for Data Integration into Knowledge Graphs

  • 提出KGpipe框架,整合信息抽取、实体匹配等工具形成端到端流程
  • 在多格式数据(RDF/JSON/文本)上测试,验证不同管道的性能差异
  • 适合需要自动化知识图谱构建的研究者与工业应用开发者

从多样化数据源构建高质量知识图谱(KGs)需融合信息抽取、数据转换、本体映射、实体匹配和数据融合等多种方法。尽管各类任务已有大量工具,但缺乏对这些方法组合成可复现、高效端到端管道的支持。本文提出KGpipe框架,用于定义与执行集成管道,可结合现有工具或大语言模型(LLM)功能。为评估不同管道及其生成的KG,我们设计了一个基准测试,将异构数据(包括RDF、JSON、文本)集成到种子知识图谱中。通过运行并对比多个管道,涵盖相同或不同格式的数据源,使用选定的性能与质量指标,验证了KGpipe的灵活性与有效性。

原文摘要 · Abstract (English)

Building high-quality knowledge graphs (KGs) from diverse sources requires combining methods for information extraction, data transformation, ontology mapping, entity matching, and data fusion. Numerous methods and tools exist for each of these tasks, but support for combining them into reproducible and effective end-to-end pipelines is still lacking. We present a new framework, KGpipe for defining and executing integration pipelines that can combine existing tools or LLM (Large Language Model) functionality. To evaluate different pipelines and the resulting KGs, we propose a benchmark to integrate heterogeneous data of different formats (RDF, JSON, text) into a seed KG. We demonstrate the flexibility of KGpipe by running and comparatively evaluating several pipelines integrating sources of the same or different formats using selected performance and quality metrics.

知识图谱数据集成LLM管道框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。