用机器学习预测糖尿病风险,神经网络准确率达78.57%。
Diabetes Prediction and Management Using Machine Learning Approaches
- 基于临床数据训练多种机器学习模型,比较其预测效果。
- 神经网络准确率最高达78.57%,随机森林次之为76.30%。
- 结果可辅助早期筛查,适合医疗决策支持系统参考。
糖尿病已成为全球重大健康问题,尤其在多国病例持续上升背景下,亟需加强早期检测与主动管理以减少严重并发症。近年来,机器学习算法在糖尿病风险预测方面展现出良好潜力。本研究基于Pima Indians Diabetes Database中的768个样本,评估统计与非统计机器学习方法在糖尿病风险分类中的表现,重点关注年龄、身体质量指数(BMI)和血糖水平等关键临床特征。实验对比了逻辑回归、决策树、随机森林、K近邻、朴素贝叶斯、支持向量机、梯度提升及神经网络等多种算法的准确性与有效性。结果显示,神经网络模型预测准确率最高,达78.57%,随机森林紧随其后,准确率为76.30%。研究证明机器学习不仅高效,还可作为数据驱动的早期筛查工具,为高风险人群提供预警信息,助力长期及时干预,降低糖尿病对医疗系统的负担。
原文摘要 · Abstract (English)
Diabetes has emerged as a significant global health issue, especially with the increasing number of cases in many countries. This trend Underlines the need for a greater emphasis on early detection and proactive management to avert or mitigate the severe health complications of this disease. Over recent years, machine learning algorithms have shown promising potential in predicting diabetes risk and are beneficial for practitioners. Objective: This study highlights the prediction capabilities of statistical and non-statistical machine learning methods over Diabetes risk classification in 768 samples from the Pima Indians Diabetes Database. It consists of the significant demographic and clinical features of age, body mass index (BMI) and blood glucose levels that greatly depend on the vulnerability against Diabetes. The experimentation assesses the various types of machine learning algorithms in terms of accuracy and effectiveness regarding diabetes prediction. These algorithms include Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbors, Naive Bayes, Support Vector Machine, Gradient Boosting and Neural Network Models. The results show that the Neural Network algorithm gained the highest predictive accuracy with 78,57 %, and then the Random Forest algorithm had the second position with 76,30 % accuracy. These findings show that machine learning techniques are not just highly effective. Still, they also can potentially act as early screening tools in predicting Diabetes within a data-driven fashion with valuable information on who is more likely to get affected. In addition, this study can help to realize the potential of machine learning for timely intervention over the longer term, which is a step towards reducing health outcomes and disease burden attributable to Diabetes on healthcare systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。