arXiv:2505.15844q-bio.QMcs.LG2025-05被引 1

用十项常见数据构建高精度脑卒中预测模型,准确率达97.2%。

Advancing Tabular Stroke Modelling Through a Novel Hybrid Architecture and Feature-Selection Synergy

  • 融合多种算法与特征选择,构建混合预测框架。
  • 在4981条数据上实现97.2%准确率与97.15% F1值。
  • 适合医疗数据建模、临床风险评估场景使用。

脑卒中仍是全球主要致死致残原因,但多数基于表格数据的预测模型准确率仍低于95%,限制了实际应用。本文基于公开的4,981条记录,利用十项常规人口学、生活方式和临床变量,构建了一个完全数据驱动且可解释的机器学习框架以预测脑卒中。通过详尽的探索性数据分析(EDA)理解数据结构,再经缺失值处理、异常值剔除及使用合成少数类过采样技术(SMOTE)修正类别不平衡等严格预处理。采用点双列相关与随机森林基尼重要性进行特征筛选,并对十种算法(包括树集成、梯度提升、核方法及多层神经网络)在分层五折交叉验证下进行优化。最终模型结合随机森林、XGBoost、LightGBM与支持向量分类器,以逻辑回归作为元学习器。该模型达到97.2%准确率与97.15% F1值,显著优于单一最优模型LightGBM的91.4%准确率。研究表明,严谨的数据预处理与多样化混合模型能将低成本表格数据转化为近乎临床级的卒中风险评估工具。

原文摘要 · Abstract (English)

Brain stroke remains one of the principal causes of death and disability worldwide, yet most tabular-data prediction models still hover below the 95% accuracy threshold, limiting real-world utility. Addressing this gap, the present work develops and validates a completely data-driven and interpretable machine-learning framework designed to predict strokes using ten routinely gathered demographic, lifestyle, and clinical variables sourced from a public cohort of 4,981 records. We employ a detailed exploratory data analysis (EDA) to understand the dataset's structure and distribution, followed by rigorous data preprocessing, including handling missing values, outlier removal, and class imbalance correction using Synthetic Minority Over-sampling Technique (SMOTE). To streamline feature selection, point-biserial correlation and random-forest Gini importance were utilized, and ten varied algorithms-encompassing tree ensembles, boosting, kernel methods, and a multilayer neural network-were optimized using stratified five-fold cross-validation. Their predictions based on probabilities helped us build the proposed model, which included Random Forest, XGBoost, LightGBM, and a support-vector classifier, with logistic regression acting as a meta-learner. The proposed model achieved an accuracy rate of 97.2% and an F1-score of 97.15%, indicating a significant enhancement compared to the leading individual model, LightGBM, which had an accuracy of 91.4%. Our study's findings indicate that rigorous preprocessing, coupled with a diverse hybrid model, can convert low-cost tabular data into a nearly clinical-grade stroke-risk assessment tool.

脑卒中预测表格数据混合模型特征选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。