arXiv:2608.18296cs.CYcs.AI2026-08

构建血糖预测公平性基准,发现模型在不同人群间存在显著误差差异。

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

论文配图:FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation
图 1 · 摘自论文原文
  • 基于300名患者构建跨12类人群的血糖数据集,用于公平性评估。
  • T1D患者预测误差比T2D高6 mg/dL,且所有模型均存在此偏差。
  • 建议采用分群报告取代单一整体指标,保障医疗AI公平性。

随着基于CGM的AI工具接近临床部署,其在不同患者群体中的准确性是否公平仍缺乏充分检验。为此,我们构建了FairGlucose,一个包含300名患者的CGM队列,按年龄、性别及1型/2型糖尿病共12个维度均衡分布,涵盖132,480个预测样本和81名患者记录的3,945个独特行为事件(如进食、运动、用药)。在4类模型家族中对33个模型进行2小时血糖预测基准测试,发现群体层面的外部验证可能掩盖显著的子群体差异。总体外分布指标稳定(约1.0),但子群体比率范围为0.8至1.4,其中T1D患者预测误差比T2D高6 mg/dL(p < 0.001)。该差异在所有33个模型中持续存在,表明其源于预测任务本质而非特定架构。进一步分析显示,子群体表现差距与临床难病例比例相关,且输入长度敏感性在不同人群中异质,支持个性化配置。前沿大模型在预测上比专用神经网络差1-6 mg/dL;即使在理想事件信息下,行为事件贡献也微乎其微(约0.1 mg/dL)。结果表明,仅依赖群体层面验证不足以评估数字健康AI的公平性,应将子群体拆解报告作为默认标准。

原文摘要 · Abstract (English)

As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.

血糖预测公平性评估数字健康医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。