用大模型构建中风知识图谱,提升医学文献信息提取精度。
SKG-LLM: Developing a Mathematical Model for Stroke Knowledge Graph Construction Using Large Language Models
- 结合数学建模与GPT-4,自动从文献中提取中风相关关系。
- 构建含2692个节点、5012条边的知识图谱,精确率达0.906,召回率达0.923。
- 适合医学信息挖掘、智能诊疗系统研发者参考。
本研究提出SKG-LLM,利用数学模型与大语言模型(LLMs)从卒中相关文献中构建知识图谱(KG),以提升卒中研究中知识的准确性和深度。方法中采用GPT-4进行数据预处理及嵌入提取。通过精确率(Precision)和召回率(Recall)评估模型性能。相较于Wikidata和WN18RR,SKG-LLM表现更优,尤其在精确率与召回率方面。引入专家评审后,精确率提升至0.923,召回率提升至0.918。最终构建的知识图谱包含2692个节点、5012条边,涵盖13种节点类型和24种边类型。
原文摘要 · Abstract (English)
The purpose of this study is to introduce SKG-LLM. A knowledge graph (KG) is constructed from stroke-related articles using mathematical and large language models (LLMs). SKG-LLM extracts and organizes complex relationships from the biomedical literature, using it to increase the accuracy and depth of KG in stroke research. In the proposed method, GPT-4 was used for data pre-processing, and the extraction of embeddings was also done by GPT-4 in the whole KG construction process. The performance of the proposed model was tested with two evaluation criteria: Precision and Recall. For further validation of the proposed model, GPT-4 was used. Compared with Wikidata and WN18RR, the proposed KG-LLM approach performs better, especially in precision and recall. By including GPT-4 in the preprocessing process, the SKG-LLM model achieved a precision score of 0.906 and a recall score of 0.923. Expert reviews further improved the results and increased precision to 0.923 and recall to 0.918. The knowledge graph constructed by SKG-LLM contains 2692 nodes and 5012 edges, which are 13 distinct types of nodes and 24 types of edges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。