arXiv:2503.21241cs.LGcs.AI2025-03被引 4

用临床特征提升医院死亡率预测准确率,随机森林表现最佳。

Feature-Enhanced Machine Learning for All-Cause Mortality Prediction in Healthcare Data

  • 基于临床经验提取生命体征、化验和人口统计特征
  • 随机森林模型AUC达0.94,优于其他机器学习与深度学习方法
  • 适合临床决策支持系统开发,尤其关注特征工程价值

准确的患者死亡率预测有助于风险分层,实现个性化治疗并改善预后。然而,医疗死亡率预测仍具挑战性,现有研究多聚焦特定疾病或有限预测变量。本研究利用MIMIC-III数据库,评估机器学习模型在全因院内死亡率预测中的表现,采用基于临床知识和文献的全面特征工程方法。提取的关键特征包括生命体征(如心率、血压)、实验室结果(如肌酐、葡萄糖)及人口统计信息。随机森林模型表现最优,AUC达0.94,显著优于其他机器学习与深度学习方法。这表明随机森林在处理高维、噪声大的临床数据方面具有鲁棒性,具备发展临床决策支持工具的潜力。研究强调了精心特征工程对准确预测的重要性,并讨论了临床应用前景,提出未来方向包括增强模型鲁棒性及针对特定疾病定制预测模型。

原文摘要 · Abstract (English)

Accurate patient mortality prediction enables effective risk stratification, leading to personalized treatment plans and improved patient outcomes. However, predicting mortality in healthcare remains a significant challenge, with existing studies often focusing on specific diseases or limited predictor sets. This study evaluates machine learning models for all-cause in-hospital mortality prediction using the MIMIC-III database, employing a comprehensive feature engineering approach. Guided by clinical expertise and literature, we extracted key features such as vital signs (e.g., heart rate, blood pressure), laboratory results (e.g., creatinine, glucose), and demographic information. The Random Forest model achieved the highest performance with an AUC of 0.94, significantly outperforming other machine learning and deep learning approaches. This demonstrates Random Forest's robustness in handling high-dimensional, noisy clinical data and its potential for developing effective clinical decision support tools. Our findings highlight the importance of careful feature engineering for accurate mortality prediction. We conclude by discussing implications for clinical adoption and propose future directions, including enhancing model robustness and tailoring prediction models for specific diseases.

死亡率预测随机森林特征工程临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。