用非侵入性数据实现糖尿病风险精准筛查,结果可解释。
Beyond the Blood Draw: Explainable Machine Learning for Non-Invasive Dysglycemia Risk Screening

- 基于健康调查数据训练六种机器学习模型,无需验血。
- 轻量梯度提升机模型AUC达0.820,优于现有临床评分工具。
- 揭示年龄、种族和腰高比是关键影响因素,适合社区筛查。
糖尿病前期和糖尿病全球影响大量成年人,但许多患者未被诊断。我们开发并验证了无需实验室检测的非侵入性糖尿病风险机器学习筛查模型。整合2017至2023年国家健康与营养调查(NHANES)数据(n=14,352),采用分层五折交叉验证训练六种机器学习模型,并与两种既定临床风险评分进行比较。轻量梯度提升机(LightGBM)表现最佳,受试者工作特征曲线下面积(AUC)为0.820(95%置信区间:0.806–0.835),显著优于芬兰糖尿病风险评分(0.745)和美国糖尿病协会风险测试(0.783)。SHAP分析表明年龄、种族/民族及腰高比为最具影响力的预测因子。亚组分析显示各人群表现一致(AUC:0.735–0.832)。研究证实了可解释的无实验室检测糖尿病风险筛查在社区及自我健康管理应用中的可行性。
原文摘要 · Abstract (English)
Dysglycemia, encompassing both prediabetes and diabetes, affects huge numbers of adults worldwide, yet many of them remain undiagnosed. We developed and validated machine-learning (ML) models for non-invasive screening of dysglycemia risk that require no laboratory tests. Pooling data from the National Health and Nutrition Examination Survey (NHANES) 2017--2023 (n=14,352), we trained six ML models with stratified 5-fold cross-validation and compared them with two established clinical risk scores. LightGBM achieved the highest area under the receiver operating characteristic curve (AUC=0.820, 95% CI: 0.806--0.835), outperforming the Finnish Diabetes Risk Score (0.745) and American Diabetes Association Risk Test (0.783). SHAP analysis identified age, race/ethnicity, and waist-to-height ratio as the most influential predictors. Subgroup analyses confirmed consistent performance across demographic strata (AUC: 0.735--0.832). These results demonstrate the feasibility of explainable, laboratory-free dysglycemia screening for deployment in community settings and self-tracking health applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。