arXiv:2512.02023cs.LG2025-12中稿 · presentation at th…被引 1

用优化特征和集成学习,提升糖尿病早期预测准确率。

An Improved Ensemble-Based Machine Learning Model with Feature Optimization for Early Diabetes Prediction

  • 结合SMOTE与Tomek Links处理数据不平衡,用堆叠模型融合XGBoost和KNN
  • 在25万条健康调查数据上达到94.82%准确率,ROC-AUC达0.989
  • 开发了基于React Native的可穿戴健康工具,适合临床筛查与个人监测

糖尿病是全球重大健康问题,早期干预依赖精准检测。但风险因素重叠与数据不平衡使预测困难。本研究利用2015年美国行为危险因素监测系统(BRFSS)数据集,包含约253,680条记录与22个数值型特征,评估多种监督学习方法。通过SMOTE与Tomek Links解决类别不平衡问题,并比较单个模型与集成方法(如堆叠)。随机森林、XGBoost、CatBoost与LightGBM等个体模型均取得约0.96的ROC-AUC表现。其中,以XGBoost与KNN为基础的堆叠集成模型表现最优:准确率达94.82%,ROC-AUC为0.989,PR-AUC达0.991,兼具高召回与高精度。此外,我们开发了一个基于React Native与Python Flask的移动端应用,实现糖尿病早期预警的可视化与可访问性,助力临床决策与日常健康监测。

原文摘要 · Abstract (English)

Diabetes is a serious worldwide health issue, and successful intervention depends on early detection. However, overlapping risk factors and data asymmetry make prediction difficult. To use extensive health survey data to create a machine learning framework for diabetes classification that is both accurate and comprehensible, to produce results that will aid in clinical decision-making. Using the BRFSS dataset, we assessed a number of supervised learning techniques. SMOTE and Tomek Links were used to correct class imbalance. To improve prediction performance, both individual models and ensemble techniques such as stacking were investigated. The 2015 BRFSS dataset, which includes roughly 253,680 records with 22 numerical features, is used in this study. Strong ROC-AUC performance of approximately 0.96 was attained by the individual models Random Forest, XGBoost, CatBoost, and LightGBM.The stacking ensemble with XGBoost and KNN yielded the best overall results with 94.82\% accuracy, ROC-AUC of 0.989, and PR-AUC of 0.991, indicating a favourable balance between recall and precision. In our study, we proposed and developed a React Native-based application with a Python Flask backend to support early diabetes prediction, providing users with an accessible and efficient health monitoring tool.

糖尿病预测集成学习数据平衡健康应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。