用大模型和本体构建无标注神经科学知识图谱,提升文献信息挖掘效率。
Entity-Augmented Neuroscience Knowledge Retrieval Using Ontology and Semantic Understanding Capability of LLM
- 结合大模型与本体,从无标注文献中自动构建神经科学知识图谱
- 实体抽取F1达0.84,知识图谱使52%以上问题回答准确率提升
- 适合需要高效整合神经科学文献的研究者使用
神经科学领域的研究文献蕴含海量知识。准确检索已有信息并从中发现新洞见,对推动该领域发展至关重要。然而,当知识分散于多个来源时,现有先进检索方法常难以提取所需信息。知识图谱(KG)可整合并连接多源知识,但当前神经科学领域构建KG的方法通常依赖标注数据且需领域专业知识,而获取大规模标注数据在专业领域极具挑战。本文提出新方法,利用大语言模型(LLM)、神经科学本体及文本嵌入,从无标注大规模神经科学论文语料中构建知识图谱。我们分析了由大模型识别的神经科学文本片段的语义相关性以构建图谱,并引入实体增强型信息检索算法从图谱中提取知识。多项实验评估表明,所提方法显著提升了从无标注文献中发现知识的能力。实体与关系抽取方法性能接近已有监督方法,在无标注数据上实现0.84的实体抽取F1值。基于知识图谱获取的知识,使超过52%的PubMedQA数据集问题及基于选定神经科学实体生成的问题得到更优解答。
原文摘要 · Abstract (English)
Neuroscience research publications encompass a vast wealth of knowledge. Accurately retrieving existing information and discovering new insights from this extensive literature is essential for advancing the field. However, when knowledge is dispersed across multiple sources, current state-of-the-art retrieval methods often struggle to extract the necessary information. A knowledge graph (KG) can integrate and link knowledge from multiple sources. However, existing methods for constructing KGs in neuroscience often rely on labeled data and require domain expertise. Acquiring large-scale, labeled data for a specialized area like neuroscience presents significant challenges. This work proposes novel methods for constructing KG from unlabeled large-scale neuroscience research corpus utilizing large language models (LLM), neuroscience ontology, and text embeddings. We analyze the semantic relevance of neuroscience text segments identified by LLM for building the knowledge graph. We also introduce an entity-augmented information retrieval algorithm to extract knowledge from the KG. Several experiments were conducted to evaluate the proposed approaches. The results demonstrate that our methods significantly enhance knowledge discovery from the unlabeled neuroscience research corpus. The performance of the proposed entity and relation extraction method is comparable to the existing supervised method. It achieves an F1 score of 0.84 for entity extraction from the unlabeled data. The knowledge obtained from the KG improves answers to over 52% of neuroscience questions from the PubMedQA dataset and questions generated using selected neuroscience entities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。