用大语言模型分析住院数据,提前预测重症患者谵妄风险。
DeLLiriuM: A large language model for delirium prediction in the ICU using structured EHR
- 基于结构化病历数据训练大语言模型,利用入院24小时信息预测谵妄。
- 在超10万患者、195家医院数据上验证,准确率最高达0.84(AUROC)。
- 首个基于结构化EHR的LLM谵妄预测模型,适合临床早期干预参考。
谵妄是重症监护室(ICU)中常见的急性意识障碍,影响高达31%的患者。早期识别可促进及时干预并改善预后。尽管人工智能(AI)模型在使用结构化电子健康记录(EHR)预测ICU谵妄方面已展现潜力,但多数模型未采用前沿技术,仅限单中心或小样本。大型语言模型(LLM)具有数亿至数十亿参数,可能提升预测性能。本研究提出DeLLiriuM,一种基于结构化EHR数据的新型LLM谵妄预测模型,利用患者入院首24小时数据预测其剩余ICU期间发生谵妄的概率。我们在三个大型数据库——eICU协作研究数据库、MIMIC-IV和佛罗里达大学医疗整合数据仓库——中涵盖195家医院的104,303例患者进行开发与验证。在两个外部验证集上,模型的受试者工作特征曲线下面积(AUROC)分别达到0.77(95%置信区间0.76–0.78)和0.84(95%置信区间0.83–0.85),覆盖77,543名患者及194家医院。据我们所知,DeLLiriuM是首个基于结构化EHR的LLM谵妄预测工具,性能优于采用结构化特征的深度学习基线模型,能为临床提供及时干预支持。
原文摘要 · Abstract (English)
Delirium is an acute confusional state that has been shown to affect up to 31% of patients in the intensive care unit (ICU). Early detection of this condition could lead to more timely interventions and improved health outcomes. While artificial intelligence (AI) models have shown great potential for ICU delirium prediction using structured electronic health records (EHR), most of them have not explored the use of state-of-the-art AI models, have been limited to single hospitals, or have been developed and validated on small cohorts. The use of large language models (LLM), models with hundreds of millions to billions of parameters, with structured EHR data could potentially lead to improved predictive performance. In this study, we propose DeLLiriuM, a novel LLM-based delirium prediction model using EHR data available in the first 24 hours of ICU admission to predict the probability of a patient developing delirium during the rest of their ICU admission. We develop and validate DeLLiriuM on ICU admissions from 104,303 patients pertaining to 195 hospitals across three large databases: the eICU Collaborative Research Database, the Medical Information Mart for Intensive Care (MIMIC)-IV, and the University of Florida Health's Integrated Data Repository. The performance measured by the area under the receiver operating characteristic curve (AUROC) showed that DeLLiriuM outperformed all baselines in two external validation sets, with 0.77 (95% confidence interval 0.76-0.78) and 0.84 (95% confidence interval 0.83-0.85) across 77,543 patients spanning 194 hospitals. To the best of our knowledge, DeLLiriuM is the first LLM-based delirium prediction tool for the ICU based on structured EHR data, outperforming deep learning baselines which employ structured features and can provide helpful information to clinicians for timely interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。