构建首个评估大模型生成知识图谱能力的基准测试
BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

- 设计多层级复杂度的真实场景测试集
- 发现主流大模型在复杂场景下表现受限
- 为自动化知识图谱构建提供可衡量标准
知识图谱生成是知识工程师最耗时且劳动密集的任务之一,因需识别输入数据源与本体术语间的语义等价关系。尽管声明式解决方案(如RML、SPARQL-Anything)已使该过程更具通用性,但将输入模式元素与本体术语对齐仍涉及复杂转换,需大量手动工作。随着大语言模型(LLMs)的发展,利用其能力辅助知识图谱构建成为研究热点。尽管已有研究探索使用LLMs自动化知识图谱构建,但尚无标准化框架评估其在建立数据模式与本体概念间对应关系方面的有效性。因此,本文提出BLINKG,一个旨在评估LLMs从异构数据源构建知识图谱映射能力的基准。该基准包含基于真实用例、复杂度逐步提升的一系列场景。我们对多个先进LLMs进行了广泛实验评估,发现它们已展现出有前景的解决方案,但在复杂场景中性能仍受限。通过此基准,我们可初步评估当前LLMs在知识图谱构建中的能力。此外,我们定义了实现(半)自动(基于LLM)知识图谱构建所需的一组要求,为该领域开启新的研究方向。
原文摘要 · Abstract (English)
Generating Knowledge Graphs (KGs) remains one of the most time-consuming and labor-intensive tasks for knowledge engineers, as they need to identify semantic equivalences between input data sources and ontology terms. While declarative solutions (e.g., RML, SPARQL-Anything) have helped to generalize this process, aligning input schema elements with ontology terms still involves intricate transformations and requires considerable manual effort. With the advent of Large Language Models (LLMs), there is growing interest in leveraging their capabilities to assist KG engineers. Although some studies have explored using LLMs to automate KG construction, there is still no standardized framework for assessing how effectively they establish correspondences between data schemes and ontology concepts. Therefore, in this paper, we propose BLINKG, a benchmark designed to evaluate the mapping capabilities of LLMs in constructing KGs from heterogeneous data sources. The benchmark includes a set of scenarios with increasing complexity, based on real-world use cases. We conduct an extensive experimental evaluation of several stateof-the-art LLMs using BLINK and observe that they already offer promising solutions. However, their performance remains limited in complex scenarios. Thanks to this benchmark, we can already assess the current capabilities of LLMs for KG construction. Additionally, we define a set of requirements for achieving (semi)automated (LLM-driven) KG construction, opening new research lines in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。