用大模型自动填充知识图谱,准确率达90%。
Ontology Population using LLMs
- 用模块化本体引导提示,让大模型从文本中提取三元组。
- 在真实数据上,提取准确率接近90%。
- 适合需要快速构建知识图谱的研究者参考。
知识图谱(KG)在数据整合、表示与可视化中日益重要。尽管知识图谱的构建至关重要,但其成本较高,尤其是从自然语言的非结构化文本中提取信息时,面临歧义和复杂理解等挑战。大语言模型(LLMs)在自然语言理解与内容生成方面表现优异,但存在幻觉问题,可能导致输出错误。然而,通过提示工程与微调,它们能实现接近人类水平的数据提取与结构化能力,具备快速且可扩展的优势。本研究聚焦于大模型在知识图谱构建中的有效性,以Enslaved.org Hub Ontology为案例,结果表明:当在提示中提供模块化本体作为指导时,大模型可提取约90%的三元组,接近真实标注数据。
原文摘要 · Abstract (English)
Knowledge graphs (KGs) are increasingly utilized for data integration, representation, and visualization. While KG population is critical, it is often costly, especially when data must be extracted from unstructured text in natural language, which presents challenges, such as ambiguity and complex interpretations. Large Language Models (LLMs) offer promising capabilities for such tasks, excelling in natural language understanding and content generation. However, their tendency to ``hallucinate'' can produce inaccurate outputs. Despite these limitations, LLMs offer rapid and scalable processing of natural language data, and with prompt engineering and fine-tuning, they can approximate human-level performance in extracting and structuring data for KGs. This study investigates LLM effectiveness for the KG population, focusing on the Enslaved.org Hub Ontology. In this paper, we report that compared to the ground truth, LLM's can extract ~90% of triples, when provided a modular ontology as guidance in the prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。