arXiv:2504.00053cs.CLcs.AI2025-04被引 12

用大模型从病历中自动识别三种常见病,准确率高且趋势一致。

Integrating Large Language Models with Human Expertise for Disease Detection in Electronic Health Records

  • 用提示词引导大模型解析临床笔记,结合诊断指南自动判别疾病
  • 检测心梗、糖尿病和高血压的敏感度超88%,对心梗和高血压特异度提升显著
  • 适合医疗数据挖掘、疾病监测研究者使用,尤其在标注资源有限时

电子健康记录(EHR)可补充基于行政数据的疾病监测与医疗评估。但从中定义疾病需大量人工标注。本研究开发了一种基于先进大语言模型(LLM)的高效策略,从临床笔记中识别多种疾病。研究将2015年阿尔伯塔省的心脏病队列与EHR系统关联,构建了利用生成式大模型通过特定诊断、治疗管理及临床指南提示词分析笔记的流程,用于检测急性心肌梗死(AMI)、糖尿病和高血压。共纳入3,088名患者,551,095条临床笔记,患病率分别为55.4%、27.7%、65.9%。性能以临床医生验证诊断为金标准,对比国际疾病分类(ICD)编码方法。结果显示:AMI敏感度88%、特异度63%、阳性预测值(PPV)77%;糖尿病敏感度91%、特异度86%、PPV 71%;高血压敏感度94%、特异度32%、PPV 72%。相比ICD编码,该方法在所有疾病上均提升敏感度与阴性预测值,且每月病例趋势与金标准一致。

原文摘要 · Abstract (English)

Objective: Electronic health records (EHR) are widely available to complement administrative data-based disease surveillance and healthcare performance evaluation. Defining conditions from EHR is labour-intensive and requires extensive manual labelling of disease outcomes. This study developed an efficient strategy based on advanced large language models to identify multiple conditions from EHR clinical notes. Methods: We linked a cardiac registry cohort in 2015 with an EHR system in Alberta, Canada. We developed a pipeline that leveraged a generative large language model (LLM) to analyze, understand, and interpret EHR notes by prompts based on specific diagnosis, treatment management, and clinical guidelines. The pipeline was applied to detect acute myocardial infarction (AMI), diabetes, and hypertension. The performance was compared against clinician-validated diagnoses as the reference standard and widely adopted International Classification of Diseases (ICD) codes-based methods. Results: The study cohort accounted for 3,088 patients and 551,095 clinical notes. The prevalence was 55.4%, 27.7%, 65.9% and for AMI, diabetes, and hypertension, respectively. The performance of the LLM-based pipeline for detecting conditions varied: AMI had 88% sensitivity, 63% specificity, and 77% positive predictive value (PPV); diabetes had 91% sensitivity, 86% specificity, and 71% PPV; and hypertension had 94% sensitivity, 32% specificity, and 72% PPV. Compared with ICD codes, the LLM-based method demonstrated improved sensitivity and negative predictive value across all conditions. The monthly percentage trends from the detected cases by LLM and reference standard showed consistent patterns.

疾病检测大模型电子病历临床挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。