一套三阶段框架,同时实现糖尿病检测、分型与血糖认知关联分析。
A Unified Three-Stage Machine Learning Framework for Diabetes Detection, Subtype Discrimination, and Cognitive-Metabolic Hypothesis Testing

- 分三阶段:检测、无监督分型、认知-代谢关联验证。
- 血糖、胰岛素和年龄是关键预测因子,分型效果良好(轮廓系数≈0.116)。
- 发现血糖控制越好,认知功能越强(ρ_s=0.208,p<0.0001),适合临床研究者参考。
全球超过5370万成人受糖尿病影响,预防医学面临重大挑战。现有机器学习研究多将糖尿病预测视为二分类问题,而分型分析与血糖-认知关联仍较少被探索。本文提出一个可复现的三阶段机器学习框架,用于糖尿病检测、分型聚类及代谢-认知关联分析。第一阶段在NCSU糖尿病数据集上,采用分层五折交叉验证,对比五种监督分类器与堆叠集成模型,评估指标包括ROC-AUC、平衡准确率、召回率和F1分数。SVM-RBF与逻辑回归取得最高ROC-AUC(0.825±0.026),随机森林准确率最高(0.762±0.030)。SHAP可解释性分析识别出葡萄糖、体重指数(BMI)和年龄为关键预测生物标志物。第二阶段使用轮廓系数验证的K-Means聚类(k=2,轮廓系数≈0.116),基于葡萄糖、胰岛素和年龄对确诊糖尿病患者进行聚类,恢复出具有临床意义的分型分区,无需真实分型标签。第三阶段对俄亥俄纵向认知数据集(n=373)进行统计分析,发现血糖控制与认知功能存在显著正相关(ρ_s=0.208,p=5.29×10⁻⁵),该结果通过霍尔姆校正。研究支持构建可解释、统计严谨的机器学习流程,推动糖尿病可复现分析与分型导向探索。
原文摘要 · Abstract (English)
Diabetes mellitus affects over 537 million adults worldwide and remains a major challenge in preventive healthcare. Existing machine-learning studies primarily formulate diabetes prediction as a binary classification problem, while subtype-oriented analysis and glycaemic-cognitive associations remain comparatively underexplored. We present a reproducible three-stage machine learning framework for diabetes detection, subtype-oriented clustering, and metabolic-cognitive association analysis. In Stage 1, five supervised classifiers together with a stacking ensemble are benchmarked on the NCSU Diabetes Dataset using stratified five-fold cross-validation and evaluation metrics including ROC-AUC, balanced accuracy, recall, and F1-score. SVM-RBF and Logistic Regression achieve the highest ROC-AUC ($0.825 \pm 0.026$), while Random Forest achieves the highest accuracy ($0.762 \pm 0.030$). SHAP explainability identifies Glucose, BMI, and Age as the dominant predictive biomarkers. In Stage 2, silhouette-validated K-Means clustering ($k=2$, silhouette $\approx 0.116$) is applied to confirmed diabetic cases using Glucose, Insulin, and Age, recovering clinically plausible subtype-oriented partitions without requiring ground-truth subtype labels. In Stage 3, statistical analysis of the Ohio Longitudinal Cognitive Dataset ($n=373$) reveals a significant positive association between glycaemic control and cognitive function ($ρ_s = 0.208$, $p = 5.29 \times 10^{-5}$), which survives Holm correction. The findings support the utility of statistically grounded and interpretable ML pipelines for reproducible diabetes analytics and subtype-aware exploratory analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。