构建首个动态知识图谱更新基准,支持文本知识与图谱结构融合。
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
- 基于维基数据快照与维基百科段落,标注145万次知识图谱更新操作。
- 覆盖2019至2025年7个年度快照,共23.3万段文本匹配更新动作。
- 为知识图谱动态演化研究提供可复现的基准数据集,适合知识管理方向研究者。
知识图谱(KG)是包含实体及其关系的结构化知识库。本文研究如何自动响应非结构化文本中不断演化的知识,对知识图谱进行时序更新。该任务需根据当前知识图谱状态与从文本中提取的信息,识别多样化的更新操作,不同于传统信息抽取独立于图谱状态的流程。为此,我们提出构建一个数据集:包含从2019到2025年共7个年度的维基数据(Wikidata)快照,以及对应维基百科段落和其引发的知识图谱更新操作。该数据集包含23.3万条维基百科段落,关联总计145万次知识图谱编辑操作。实验结果揭示了在将文本表达的知识与现有图谱结构融合方面存在关键挑战。本数据集为未来研究提供了重要基准。数据集与模型实现均已公开。
原文摘要 · Abstract (English)
Knowledge Graphs (KGs) are structured knowledge repositories containing entities and relations between them. In this paper, we study the problem of automatically updating KGs over time in response to evolving knowledge in unstructured textual sources. Addressing this problem requires identifying a wide range of update operations based on the state of an existing KG at a given time and the information extracted from text. This contrasts with traditional information extraction pipelines, which extract knowledge from text independently of the current state of a KG. To address this challenge, we propose a method for construction of a dataset consisting of Wikidata KG snapshots over time and Wikipedia passages paired with the corresponding edit operations that they induce in a particular KG snapshot. The resulting dataset comprises 233K Wikipedia passages aligned with a total of 1.45 million KG edits over 7 different yearly snapshots of Wikidata from 2019 to 2025. Our experimental results highlight key challenges in updating KG snapshots based on emerging textual knowledge, particularly in integrating knowledge expressed in text with the existing KG structure. These findings position the dataset as a valuable benchmark for future research. Our dataset and model implementations are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。