构建首个核聚变能源知识图谱,提升专业信息提取与问答效率
Automated Construction of a Knowledge Graph of Nuclear Fusion Energy for Effective Elicitation and Retrieval of Information
- 基于大模型自动识别实体并解决歧义,构建核聚变领域知识图谱
- 在复杂多跳查询上实现精准回答,准确率优于传统方法
- 适合科研人员、政策制定者快速获取核聚变领域核心知识
本文提出一种多阶段自动化方法,用于从大规模文档中构建领域知识图谱。以高度专业且异质的核聚变能源领域为应用对象,首次构建了该领域的知识图谱,作为测试管道关键能力的理想基准。该方法利用预训练大语言模型实现自动命名实体识别与实体消歧,并通过对比齐普夫定律验证其在自然语言中的表现。此外,开发了基于知识图谱的检索增强生成系统,结合多种提示策略,支持对自然语言问题(包括需跨实体推理的多跳问题)提供上下文相关答案。
原文摘要 · Abstract (English)
In this document, we discuss a multi-step approach to automated construction of a knowledge graph, for structuring and representing domain-specific knowledge from large document corpora. We apply our method to build the first knowledge graph of nuclear fusion energy, a highly specialized field characterized by vast scope and heterogeneity. This is an ideal benchmark to test the key features of our pipeline, including automatic named entity recognition and entity resolution. We show how pre-trained large language models can be used to address these challenges and we evaluate their performance against Zipf's law, which characterizes human natural language. Additionally, we develop a knowledge-graph retrieval-augmented generation system that uses multiple prompts with large language models to provide contextually relevant answers to natural-language queries, including complex multi-hop questions requiring reasoning across interconnected entities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。