用大模型自动生成知识图谱模式,提升数据质量与可读性。
Schema Generation for Large Knowledge Graphs Using Large Language Models
- 结合知识图谱的局部与全局信息,用大模型生成形状表达式模式
- 在YAGO和Wikidata数据集上生成高质量的ShEx模式,准确率达78%
- 适合知识工程、语义网研究者及自动化数据建模场景
模式在语义网和自然语言处理中对保障数据质量和提升可用性至关重要。传统模式构建依赖知识工程师和领域专家,成本高昂。本文利用大语言模型(LLMs)在本体工程中的强大能力,探索基于LLM的模式生成方法。为弥补资源不足,我们构建了两个新数据集:YAGO Schema 和 Wikidata EntitySchema,并提出新的评估指标。基于知识图谱的局部与全局信息,提出的LLM流水线可生成符合形状表达式(ShEx)规范的模式。实验表明,该方法能有效生成高质量的ShEx模式,为大规模知识图谱的可扩展自动化模式生成开辟道路。此外,我们的基准挑战了结构化生成任务,推动了大模型在语法丰富的形式化语言上的边界。
原文摘要 · Abstract (English)
Schemas play a vital role in ensuring data quality and supporting usability in the Semantic Web and natural language processing. Traditionally, their creation demands substantial involvement from knowledge engineers and domain experts. Leveraging the impressive capabilities of large language models (LLMs) in tasks like ontology engineering, we explore schema generation using LLMs. To bridge the resource gap, we introduce two datasets: YAGO Schema and Wikidata EntitySchema, along with novel evaluation metrics. The LLM-based pipelines utilize local and global information from knowledge graphs (KGs) to generate schemas in Shape Expressions (ShEx). Experiments demonstrate LLMs' strong potential in producing high-quality ShEx schemas, paving the way for scalable, automated schema generation for large KGs. Furthermore, our benchmark introduces a new challenge for structured generation, pushing the limits of LLMs on syntactically rich formalisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。