arXiv:2412.15256cs.CLcs.AI2024-12被引 5

用大模型从病历中提取疾病知识,实现精准患者搜索与新病种发现。

Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search

  • 用大模型直接从电子病历中提取疾病实体,无需依赖固定分类体系。
  • 在3360万患者数据中成功定位Dravet综合征患者,准确率高且可追溯。
  • 首次实现无先验知识的罕见病BPAN患者发现,适合临床研究与医学探索者。

构建和维护知识图谱可加速真实世界数据中的疾病发现与分析。尽管疾病本体(如SNOMED-CT、ICD10、CPT)有助于生物数据标注,但其编码类别可能无法捕捉患者病情细节或罕见病特征。多源数据中的疾病定义差异使本体映射与疾病聚类复杂化。本文提出使用大语言模型(LLM)提取技术构建患者知识图谱,支持自然语言方式的数据提取,而非依赖僵化的本体层级。该方法将抽取实体映射至现有本体(MeSH、SNOMED-CT、RxNORM、HPO)以实现语义锚定。基于包含3360万患者的门诊电子健康记录(EHR)数据库,我们以2020年10月获得ICD10认定的Dravet综合征为例,演示了患者搜索与个体化知识图谱构建。利用已确认的Dravet综合征ICD10编码作为真值,采用基于LLM的实体抽取方法对患者进行本体化表征。随后,将该方法应用于识别β-螺旋桨蛋白相关神经退行性疾病(BPAN)患者,实现了无真值条件下的真实世界新病种发现。

原文摘要 · Abstract (English)

Creation and curation of knowledge graphs can accelerate disease discovery and analysis in real-world data. While disease ontologies aid in biological data annotation, codified categories (SNOMED-CT, ICD10, CPT) may not capture patient condition nuances or rare diseases. Multiple disease definitions across data sources complicate ontology mapping and disease clustering. We propose creating patient knowledge graphs using large language model extraction techniques, allowing data extraction via natural language rather than rigid ontological hierarchies. Our method maps to existing ontologies (MeSH, SNOMED-CT, RxNORM, HPO) to ground extracted entities. Using a large ambulatory care EHR database with 33.6M patients, we demonstrate our method through the patient search for Dravet syndrome, which received ICD10 recognition in October 2020. We describe our construction of patient-specific knowledge graphs and symptom-based patient searches. Using confirmed Dravet syndrome ICD10 codes as ground truth, we employ LLM-based entity extraction to characterize patients in grounded ontologies. We then apply this method to identify Beta-propeller protein-associated neurodegeneration (BPAN) patients, demonstrating real-world discovery where no ground truth exists.

医疗知识图谱大模型应用罕见病发现自然语言检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。