arXiv:2607.16253cs.LGcs.AI2026-07中稿 · and published at t…

评估机器学习预测糖尿病风险的公平性,发现模型对老年人和肥胖者效果差。

Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis

  • 用XGBoost模型基于8个非实验室指标训练,涵盖年龄、体重、活动等
  • 外部验证显示模型性能下降9.7%,老年群体准确率低13.5个百分点
  • 揭示算法偏见,建议按年龄分层部署,避免不公平临床应用

基于机器学习的2型糖尿病风险预测模型在内部验证中表现良好,但在真实场景中因缺乏外部测试与公平性评估而失效。本文构建多维度框架,在全国代表性人群中评估区分度、校准性、可解释性及算法公平性。使用NHANES 2015-2020数据(n=15,685)训练一个基于八项非实验室指标(年龄、性别、种族/族裔、BMI、吸烟状态、体力活动、心梗史、卒中史)的XGBoost模型。在BRFSS 2020-2022数据集(n=1,285,783)上进行外部验证,模拟现实分布偏移。内部验证区分度良好(AUC=0.794,95% CI 0.788–0.800),外部验证性能显著下降(AUC=0.717,相对降低9.7%,p<0.001)。公平性分析显示严重偏差:60岁及以上人群AUC=0.607,远低于年轻人的0.742(差异=0.135,p<0.001);肥胖者(AUC=0.698)低于正常体重者(AUC=0.735,差异=0.037,p<0.001)。性别间表现相近(男性=0.723,女性=0.712,p=0.142)。校准结果显示风险高估(Brier score=0.123)。SHAP分析指出年龄、BMI和体力活动为主要驱动因素。高风险人群反而获得最差算法表现,凸显临床部署前需采用公平性感知、年龄分层策略。

原文摘要 · Abstract (English)

Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment. We developed a multi-dimensional framework evaluating discrimination, calibration, interpretability, and algorithmic fairness on nationally representative populations. An XGBoost model was trained on NHANES 2015-2020 (n=15,685) using eight non-laboratory predictors: age, sex, race/ethnicity, BMI, smoking status, physical activity, history of heart attack, and history of stroke. External validation was performed on BRFSS 2020-2022 (n=1,285,783) under realistic distribution shift. Internal validation showed good discrimination (AUC=0.794, 95% CI 0.788-0.800), with performance loss on external validation (AUC=0.717, relative decrease: -9.7%, p<0.001). Fairness analysis revealed severe bias: elderly adults (>=60) showed AUC=0.607 vs 0.742 for young adults (difference=0.135, p<0.001); obese individuals showed AUC=0.698 vs 0.735 for normal weight (difference=0.037, p<0.001). Gender showed comparable performance (male=0.723 vs female=0.712, p=0.142). Calibration revealed risk overestimation (Brier score=0.123). SHAP analysis identified age, BMI, and physical activity as primary risk drivers. Populations with highest diabetes risk receive the worst algorithmic performance, underscoring the need for fairness-aware, age-stratified deployment strategies before clinical use.

糖尿病预测公平性评估机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。