用罗马尼亚新数据集提升脓毒症预后预测的可解释性
Explainable Machine Learning for Sepsis Outcome Prediction Using a Novel Romanian Electronic Health Record Dataset

- 基于600项检验构建可解释机器学习模型,融合临床指标
- 死亡与康复预测AUC达0.983,准确率93%
- 发现嗜酸性粒细胞减少等被忽视但关键的预警指标
我们利用罗马尼亚大型急诊医院12,286例住院记录构建的新型电子健康记录(EHR)数据集,开发并分析了用于脓毒症预后预测的可解释机器学习(ML)模型。数据集包含人口统计学信息、国际疾病分类(ICD-10)诊断编码及600类实验室检测项目。本研究旨在识别临床强预测因子,并在三项分类任务中实现顶尖性能:(1)死亡 vs. 出院,(2)死亡 vs. 恢复,(3)恢复 vs. 病情改善。训练了五种机器学习模型以捕捉复杂分布并保持临床可解释性。实验探讨了特征丰富度与患者覆盖率之间的权衡,使用了10至50个最常见实验室检测项目的子集。模型性能通过准确率和受试者工作特征曲线下面积(AUC)评估,可解释性采用SHapley Additive exPlanations(SHAP)分析。最高性能出现在死亡与恢复的对比任务中(AUC=0.983,准确率=0.93)。SHAP分析揭示多个强预测因子,如心血管共病、尿素水平、天冬氨酸氨基转移酶、血小板计数及嗜酸性粒细胞百分比。嗜酸性粒细胞减少被识别为顶级预测因子,凸显其作为未被充分利用标志物的价值,且当前评估标准尚未纳入,同时高表现力表明这些模型具备临床应用潜力。
原文摘要 · Abstract (English)
We develop and analyze explainable machine learning (ML) models for sepsis outcome prediction using a novel Electronic Health Record (EHR) dataset from 12,286 hospitalizations at a large emergency hospital in Romania. The dataset includes demographics, International Classification of Diseases (ICD-10) diagnostics, and 600 types of laboratory tests. This study aims to identify clinically strong predictors while achieving state-of-the-art results across three classification tasks: (1)deceased vs. discharged, (2)deceased vs. recovered, and (3)recovered vs. ameliorated. We trained five ML models to capture complex distributions while preserving clinical interpretability. Experiments explored the trade-off between feature richness and patient coverage, using subsets of the 10--50 most frequent laboratory tests. Model performance was evaluated using accuracy and area under the curve (AUC), and explainability was assessed using SHapley Additive exPlanations (SHAP). The highest performance was obtained for the deceased vs. recovered case study (AUC=0.983, accuracy=0.93). SHAP analysis identified several strong predictors such as cardiovascular comorbidities, urea levels, aspartate aminotransferase, platelet count, and eosinophil percentage. Eosinopenia emerged as a top predictor, highlighting its value as an underutilized marker that is not included in current assessment standards, while the high performance suggests the applicability of these models in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。