arXiv:2510.14997q-bio.OTcs.AI2025-10被引 1

用机器学习提升糖尿病患者心肾疾病的早期预警能力。

Evaluation and Implementation of Machine Learning Algorithms to Predict Early Detection of Kidney and Heart Disease in Diabetic Patients

  • 融合统计分析与随机森林等算法,筛选关键预测因子。
  • 随机森林对肾病预测准确率最高,显著优于单一模型。
  • 适合临床风险分层,助力糖尿病并发症早期干预。

心血管疾病和慢性肾病是糖尿病的主要并发症,导致高发病率和死亡率。早期检测至关重要,但传统诊断标志物在初期敏感性不足。本研究结合传统统计方法与机器学习技术,提升糖尿病患者肾病(CKD)和心脏病(CVD)的早期诊断能力。使用SPSS进行描述性和推断性统计分析,探索疾病与临床及人口学因素的关系。患者分为四组:A组(同时患CKD与CVD)、B组(仅患CKD)、C组(仅患CVD)、D组(无疾病)。统计分析显示:血清肌酐和高血压与CKD显著相关;胆固醇、甘油三酯、心肌梗死、中风和高血压与CVD显著相关。这些结果指导了机器学习模型的特征选择。采用逻辑回归、支持向量机和随机森林算法,其中随机森林在CKD预测中表现最佳。集成模型在识别高危患者方面优于单个分类器。SPSS结果进一步验证了模型所用关键参数的显著性。尽管存在可解释性和类别不平衡等挑战,该混合统计-机器学习框架相比传统方法,在糖尿病并发症的早期检测与风险分层方面展现出显著优势。

原文摘要 · Abstract (English)

Cardiovascular disease and chronic kidney disease are major complications of diabetes, leading to high morbidity and mortality. Early detection of these conditions is critical, yet traditional diagnostic markers often lack sensitivity in the initial stages. This study integrates conventional statistical methods with machine learning approaches to improve early diagnosis of CKD and CVD in diabetic patients. Descriptive and inferential statistics were computed in SPSS to explore associations between diseases and clinical or demographic factors. Patients were categorized into four groups: Group A both CKD and CVD, Group B CKD only, Group C CVD only, and Group D no disease. Statistical analysis revealed significant correlations: Serum Creatinine and Hypertension with CKD, and Cholesterol, Triglycerides, Myocardial Infarction, Stroke, and Hypertension with CVD. These results guided the selection of predictive features for machine learning models. Logistic Regression, Support Vector Machine, and Random Forest algorithms were implemented, with Random Forest showing the highest accuracy, particularly for CKD prediction. Ensemble models outperformed single classifiers in identifying high-risk diabetic patients. SPSS results further validated the significance of the key parameters integrated into the models. While challenges such as interpretability and class imbalance remain, this hybrid statistical machine learning framework offers a promising advancement toward early detection and risk stratification of diabetic complications compared to conventional diagnostic approaches.

疾病预测机器学习糖尿病风险分层

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。