arXiv:2508.12393cs.CLcs.AI2025-08中稿 · Npj Digital Medici…被引 6

用大模型每天自动构建会随时间演化的医学知识图谱

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

  • 设计双智能体框架,每日从海量文献中提取并整合医学知识
  • 构建含156,275实体、297万三元组的全球最大医用药学知识图谱
  • 支持动态演化,提升医疗问答等任务的准确率,适合临床科研人员

医学文献的快速扩张给领域知识的规模化结构化带来挑战。知识图谱(KG)提供了解决方案,但现有构建方法缺乏泛化能力,且忽略知识的时序动态性。为此,我们提出MedKGent——一个基于大语言模型(LLM)的智能体框架,用于构建时序演化的医学知识图谱。利用1975至2023年间超过1000万篇PubMed摘要,MedKGent通过两个专用智能体每日增量构建知识图谱:提取器智能体识别知识三元组并赋信心评分,构造器智能体将三元组融入时序图谱,强化重复知识并解决冲突。最终生成的图谱包含156,275个实体和2,971,384个三元组,据我们所知是目前最大的由大模型生成的医学知识图谱。自动化与专家评估显示三元组有效性接近90%。在下游任务中,MedKGent-KG显著提升了五个大模型在七项医学问答基准上的检索增强生成性能。这些结果表明,MedKGent为医学知识表示与文献驱动的人工智能研究提供了可扩展、时序感知的基础架构。

原文摘要 · Abstract (English)

The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet current construction methods lack generalizability and ignore the temporal dynamics of evolving knowledge. To address this, we introduce MedKGent, a Large Language Model (LLM) agent framework for building temporally evolving medical KGs. Using over 10 million PubMed abstracts from 1975 to 2023, MedKGent incrementally constructs a KG daily via two specialized agents. The Extractor Agent identifies knowledge triples and assigns confidence scores, while the Constructor Agent integrates these triples into a temporal graph, reinforcing recurring knowledge and resolving conflicts. The resulting KG contains 156,275 entities and 2,971,384 triples, making it, to our knowledge, the largest LLM-derived medical KG to date. Automated and expert assessments showed triple-validity rates approaching 90%. In downstream evaluations, MedKGent-KG significantly improved retrieval-augmented generation for five LLMs across seven medical question-answering benchmarks. Together, these results position MedKGent as a scalable and temporally aware infrastructure for medical knowledge representation and literature-grounded AI research.

知识图谱医学AI大模型应用时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。