用生活方式数据预测糖尿病,机器学习模型表现良好。
A Comparative Study of Diabetes Prediction Based on Lifestyle Factors Using Machine Learning
- 基于行为风险调查数据,用决策树、KNN和逻辑回归建模。
- 逻辑回归准确率达75%,三模型召回率均超70%。
- 适合医疗预防研究者参考,可拓展特征优化与集成方法。
糖尿病是全球范围内的高发慢性病,带来重大健康与经济负担。早期预测与诊断有助于有效管理并预防并发症。本研究利用2015年行为风险因素监测系统(BRFSS)数据,基于21个生活方式与健康相关特征(如体力活动、饮食、心理健康、社会经济状况),采用决策树、K近邻(KNN)和逻辑回归三种分类模型进行糖尿病预测。使用平衡数据集训练与测试模型,评估指标包括准确率、精确率、召回率和F1分数。结果表明,决策树、KNN和逻辑回归的准确率分别为0.74、0.72和0.75,各模型在精确率与召回率上表现各异。研究验证了机器学习在糖尿病预测中的潜力,并建议未来通过特征选择与集成学习进一步提升性能。
原文摘要 · Abstract (English)
Diabetes is a prevalent chronic disease with significant health and economic burdens worldwide. Early prediction and diagnosis can aid in effective management and prevention of complications. This study explores the use of machine learning models to predict diabetes based on lifestyle factors using data from the Behavioral Risk Factor Surveillance System (BRFSS) 2015 survey. The dataset consists of 21 lifestyle and health-related features, capturing aspects such as physical activity, diet, mental health, and socioeconomic status. Three classification models, Decision Tree, K-Nearest Neighbors (KNN), and Logistic Regression, are implemented and evaluated to determine their predictive performance. The models are trained and tested using a balanced dataset, and their performances are assessed based on accuracy, precision, recall, and F1-score. The results indicate that the Decision Tree, KNN, and Logistic Regression achieve an accuracy of 0.74, 0.72, and 0.75, respectively, with varying strengths in precision and recall. The findings highlight the potential of machine learning in diabetes prediction and suggest future improvements through feature selection and ensemble learning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。