融合文本数据提升心脏骤停患者死亡率预测准确率
Enhancing mortality prediction in cardiac arrest ICU patients through meta-modeling of structured clinical data from MIMIC-IV
- 用结构化数据+文本报告构建混合模型,结合特征选择与逻辑回归
- 加入文本信息后AUC达0.918,比纯结构化数据提升22%
- 模型在多种临床阈值下均表现更优,适合重症监护场景使用
准确预测重症监护室(ICU)患者住院死亡率对及时干预和资源分配至关重要。本研究基于MIMIC-IV数据库,构建并评估了融合结构化临床数据与非结构化文本信息(出院小结、影像报告)的机器学习模型。采用LASSO和XGBoost进行特征选择,再通过多变量逻辑回归训练前几项关键特征。利用TF-IDF与BERT嵌入表示文本特征,显著提升预测性能。最终模型在结合结构化与文本输入时达到AUC 0.918,较仅用结构化数据(AUC 0.753)相对提升22%。决策曲线分析显示,在0.2至0.8的阈值范围内,模型标准化净收益更优,证实其临床实用性。结果表明,非结构化临床笔记具有额外预后价值,支持其融入可解释的风险预测模型中。
原文摘要 · Abstract (English)
Accurate early prediction of in-hospital mortality in intensive care units (ICUs) is essential for timely clinical intervention and efficient resource allocation. This study develops and evaluates machine learning models that integrate both structured clinical data and unstructured textual information, specifically discharge summaries and radiology reports, from the MIMIC-IV database. We used LASSO and XGBoost for feature selection, followed by a multivariate logistic regression trained on the top features identified by both models. Incorporating textual features using TF-IDF and BERT embeddings significantly improved predictive performance. The final logistic regression model, which combined structured and textual input, achieved an AUC of 0.918, compared to 0.753 when using structured data alone, a relative improvement 22%. The analysis of the decision curve demonstrated a superior standardized net benefit in a wide range of threshold probabilities (0.2-0.8), confirming the clinical utility of the model. These results underscore the added prognostic value of unstructured clinical notes and support their integration into interpretable feature-driven risk prediction models for ICU patients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。