arXiv:2510.26188cs.LGcs.AI2025-10

用机器学习预测住院患者再入院风险,助力降本提质

Predicting All-Cause Hospital Readmissions from Medical Claims Data of Hospitalised Patients

  • 基于医保数据,用主成分分析降维后建模
  • 随机森林表现最优,AUC领先其他模型
  • 可识别高风险人群,适合医疗管理者参考

降低可预防的再入院率是支付方、医疗机构和政策制定者关注的全国性重点。再入院率被用作衡量医院医疗质量的指标。本研究采用逻辑回归、随机森林和支持向量机等机器学习方法,分析住院患者医保数据,识别影响全因再入院的关键人口统计与医疗因素。由于医保数据维度高,使用主成分分析进行降维,并据此构建回归模型。通过曲线下面积(AUC)评估模型性能,随机森林表现最佳,其次为逻辑回归和支撑向量机。这些模型可帮助识别导致再入院的关键因素,定位高风险患者,从而降低再入院率,减少医疗成本,提升医疗质量。

原文摘要 · Abstract (English)

Reducing preventable hospital readmissions is a national priority for payers, providers, and policymakers seeking to improve health care and lower costs. The rate of readmission is being used as a benchmark to determine the quality of healthcare provided by the hospitals. In thisproject, we have used machine learning techniques like Logistic Regression, Random Forest and Support Vector Machines to analyze the health claims data and identify demographic and medical factors that play a crucial role in predicting all-cause readmissions. As the health claims data is high dimensional, we have used Principal Component Analysis as a dimension reduction technique and used the results for building regression models. We compared and evaluated these models based on the Area Under Curve (AUC) metric. Random Forest model gave the highest performance followed by Logistic Regression and Support Vector Machine models. These models can be used to identify the crucial factors causing readmissions and help identify patients to focus on to reduce the chances of readmission, ultimately bringing down the cost and increasing the quality of healthcare provided to the patients.

再入院预测机器学习医保数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。