用AI整合病历文本与结构化数据,精准预测重症患者死亡率和资源消耗。
Prediction of mortality and resource utilization in critical care: a deep learning approach using multimodal electronic health records with natural language processing techniques
- 融合文本与结构化数据的深度学习框架,引入医学提示词增强理解。
- 死亡率预测准确率提升1.6%,手术时长预测误差降低11.0%。
- 对数据缺失或错误有强鲁棒性,适合临床真实场景应用。
背景:从电子健康记录(EHR)中预测重症患者死亡率和资源使用情况,对优化治疗结果和控制成本至关重要。现有方法多依赖结构化数据,忽视自由文本病历中的临床信息,且未充分挖掘结构化数据中的文本潜力。本研究提出并评估了一种基于自然语言处理的深度学习框架,整合多模态EHR以预测重症监护中的死亡率和资源利用。方法:基于两个真实世界EHR数据集,我们在三个临床任务上对比了主流方法,并对模型中三个关键组件——医学提示词、自由文本、预训练句向量编码器——进行了消融实验,同时评估了模型在结构化数据受损情况下的鲁棒性。结果:在两个真实数据集上的三任务实验表明,该模型在死亡率预测中BACC/AUROC分别提升1.6%/0.8%,住院时长预测中RMSE/MAE提升0.5%/2.2%,手术时长估计中RMSE/MAE降低10.9%/11.0%,在不同数据损坏率下均优于其他基线模型。结论:该框架是预测重症患者死亡率与资源使用的有效且高精度方法;提示学习结合Transformer编码器在多模态EHR分析中表现优异;模型在高数据损坏情况下仍具强鲁棒性。
原文摘要 · Abstract (English)
Background Predicting mortality and resource utilization from electronic health records (EHRs) is challenging yet crucial for optimizing patient outcomes and managing costs in intensive care unit (ICU). Existing approaches predominantly focus on structured EHRs, often ignoring the valuable clinical insights in free-text notes. Additionally, the potential of textual information within structured data is not fully leveraged. This study aimed to introduce and assess a deep learning framework using natural language processing techniques that integrates multimodal EHRs to predict mortality and resource utilization in critical care settings. Methods Utilizing two real-world EHR datasets, we developed and evaluated our model on three clinical tasks with leading existing methods. We also performed an ablation study on three key components in our framework: medical prompts, free-texts, and pre-trained sentence encoder. Furthermore, we assessed the model's robustness against the corruption in structured EHRs. Results Our experiments on two real-world datasets across three clinical tasks showed that our proposed model improved performance metrics by 1.6\%/0.8\% on BACC/AUROC for mortality prediction, 0.5%/2.2% on RMSE/MAE for LOS prediction, 10.9%/11.0% on RMSE/MAE for surgical duration estimation compared to the best existing methods. It consistently demonstrated superior performance compared to other baselines across three tasks at different corruption rates. Conclusions The proposed framework is an effective and accurate deep learning approach for predicting mortality and resource utilization in critical care. The study also highlights the success of using prompt learning with a transformer encoder in analyzing multimodal EHRs. Importantly, the model showed strong resilience to data corruption within structured data, especially at high corruption levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。