arXiv:2508.03438cs.AI2025-08被引 1

用增强三元组构建医学知识图谱,提升信息提取准确性与可读性。

Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction

  • 通过大模型分步解析文献,提取语义命题并生成带上下文的四元组。
  • 增强后三元组生成语句与原文相似度达0.874,上下文引入显著提升表现。
  • 适用于医疗科研人员快速获取整合知识,也具跨领域扩展潜力。

公开医学数据的快速增长加剧了文献量与实际应用间的鸿沟,使临床医生和研究人员难以系统掌握最新知识。本文提出一种基于大语言模型(LLM)代理的流程,将44篇PubMed摘要分解为语义命题,并从中提取知识图谱三元组。通过融合开放域与基于本体的信息抽取方法,三元组被增强以包含本体类别。此外,在抽取中引入上下文变量,使三元组升级为可独立理解的“四元组”。通过对比由增强三元组生成的自然语言句子与原始命题,平均余弦相似度达到0.874。与普通三元组相比,加入上下文后生成句子的相似度明显提升。研究还探索了大模型推断新关系及连接知识图谱簇的能力。该方法有望为医疗从业者提供实时、集中、可持续的知识源,也为其他领域提供可借鉴的技术路径。

原文摘要 · Abstract (English)

The rapid expansion of publicly-available medical data presents a challenge for clinicians and researchers alike, increasing the gap between the volume of scientific literature and its applications. The steady growth of studies and findings overwhelms medical professionals at large, hindering their ability to systematically review and understand the latest knowledge. This paper presents an approach to information extraction and automatic knowledge graph (KG) generation to identify and connect biomedical knowledge. Through a pipeline of large language model (LLM) agents, the system decomposes 44 PubMed abstracts into semantically meaningful proposition sentences and extracts KG triples from these sentences. The triples are enhanced using a combination of open domain and ontology-based information extraction methodologies to incorporate ontological categories. On top of this, a context variable is included during extraction to allow the triple to stand on its own - thereby becoming `quadruples'. The extraction accuracy of the LLM is validated by comparing natural language sentences generated from the enhanced triples to the original propositions, achieving an average cosine similarity of 0.874. The similarity for generated sentences of enhanced triples were compared with generated sentences of ordinary triples showing an increase as a result of the context variable. Furthermore, this research explores the ability for LLMs to infer new relationships and connect clusters in the knowledge base of the knowledge graph. This approach leads the way to provide medical practitioners with a centralised, updated in real-time, and sustainable knowledge source, and may be the foundation of similar gains in a wide variety of fields.

知识图谱医学AI大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。