用外部医学知识增强大模型,提升病历数据对住院风险的预测能力
Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling
- 用微调的大模型+医学知识图谱从病历中提取关键信息
- 在读卡率和院内死亡预测上分别达0.84和0.92的AUC
- 适合需要高精度风险评估的临床决策支持场景
基于电子健康记录(EHR)的临床结果准确预测对于早期干预、资源高效分配和改善患者护理至关重要。EHR包含结构化数据和非结构化临床笔记,蕴含丰富的情境信息。本文提出一种统一框架,通过两阶段架构融合多模态数据进行临床风险预测:第一阶段利用微调的大语言模型(LLM)从临床笔记中提取任务相关关键信息,并通过图检索方式引入来自PubMed等医学语料库的外部领域知识,增强模型理解;第二阶段将非结构化表示与结构化特征融合生成最终预测。该方法适用于多种临床任务,在30天再入院和院内死亡预测上表现优异,分别达到0.84和0.92的AUC,即使在正样本率仅约4%至13%的严重不平衡数据集上仍优于所有现有基线和临床评分系统。据我们所知,这是首个将基于图引导的知识检索与结构化数据结合用于临床预测的框架。
原文摘要 · Abstract (English)
Accurate prediction of clinical outcomes using Electronic Health Records (EHRs) is critical for early intervention, efficient resource allocation, and improved patient care. EHRs contain multimodal data, including both structured data and unstructured clinical notes that provide rich, context-specific information. In this work, we introduce a unified framework that seamlessly integrates these diverse modalities, leveraging all relevant available information through a two-stage architecture for clinical risk prediction. In the first stage, a fine-tuned Large Language Model (LLM) extracts crucial, task-relevant information from clinical notes, which is enhanced by graph-based retrieval of external domain knowledge from sources such as a medical corpus like PubMed, grounding the LLM's understanding. The second stage combines both unstructured representations and features derived from the structured data to generate the final predictions. This approach supports a wide range of clinical tasks. Here, we demonstrate its effectiveness on 30-day readmission and in-hospital mortality prediction. Experimental results show that our framework achieves strong performance, with AUC scores of $0.84$ and $0.92$, respectively, despite these tasks involving severely imbalanced datasets, with positive rates ranging from approximately $4\%$ to $13\%$. Moreover, it outperforms all existing baselines and clinical practices, including established risk scoring systems. To the best of our knowledge, this is one of the first frameworks for healthcare prediction which enhances the power of an LLM-based graph-guided knowledge retrieval method by combining it with structured data for improved clinical outcome prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。