arXiv:2505.09812cs.LG2025-05被引 7

用机器学习预测中风风险,提升早期筛查可靠性

Comparative Analysis of Stroke Prediction Models Using Machine Learning

  • 对比逻辑回归、随机森林等模型在中风预测中的表现
  • 模型准确率高但敏感度仍不足,影响临床应用
  • 识别关键预测特征,助力模型优化与可解释性

中风是全球第二大死因和第三大致残原因。本研究基于中风预测数据集,利用人口统计、临床及生活方式数据,评估了多种机器学习算法在中风风险预测中的效果。针对类别不平衡和缺失数据等方法论挑战,我们比较了逻辑回归、随机森林和XGBoost等模型。结果表明,尽管各模型均达到较高准确率,但敏感度仍是制约其在真实临床场景中应用的关键瓶颈。同时,研究识别出最具影响力的预测特征,并提出改进策略。这些发现有助于构建更可靠、可解释的中风风险早期评估模型。

原文摘要 · Abstract (English)

Stroke remains one of the most critical global health challenges, ranking as the second leading cause of death and the third leading cause of disability worldwide. This study explores the effectiveness of machine learning algorithms in predicting stroke risk using demographic, clinical, and lifestyle data from the Stroke Prediction Dataset. By addressing key methodological challenges such as class imbalance and missing data, we evaluated the performance of multiple models, including Logistic Regression, Random Forest, and XGBoost. Our results demonstrate that while these models achieve high accuracy, sensitivity remains a limiting factor for real-world clinical applications. In addition, we identify the most influential predictive features and propose strategies to improve machine learning-based stroke prediction. These findings contribute to the development of more reliable and interpretable models for the early assessment of stroke risk.

中风预测机器学习医疗健康可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。