arXiv:2608.01820cs.HCcs.LG2026-08

PPG测血糖模型在严格评估下失效,临床指标掩盖了模型失败。

Reassessing the Feasibility of PPG-Based Non-Invasive Blood Glucose Level Estimation

论文配图:Reassessing the Feasibility of PPG-Based Non-Invasive Blood Glucose Level Estimation
图 1 · 摘自论文原文
  • 构建可复现的评估流水线,采用三种渐严的数据划分策略
  • 多数模型在参与者级划分下R²接近零,表现不如均值预测
  • 90%预测在临床可接受区,但模型实际无效,需警惕指标误导

从光电容积脉搏波描记(PPG)无创估算血糖水平具有可穿戴健康监测前景,但各研究结果难以比较,原因在于数据集不一致、数据泄露和评价标准不统一。本文提出首个可复现、可扩展的评估流程,基于公开数据集,在三种逐级严格的划分协议下重新评估五种代表性PPG-BGL方法:随机窗口级、参与者感知型、留部分参与者出训练(LSPO)。模型在随机划分下表现良好,但在参与者感知与LSPO评估中大幅退化,几乎所有模型的R²接近零或为负,与均值预测基线相当。关键发现:无论模型或划分方式,超过90%的预测位于临床可接受区域(Clarke误差网格A+B),包括基线。这揭示根本矛盾:临床区域指标系统性掩盖了该领域模型的真实失败。研究证明,随机训练测试划分会因样本级数据泄露而严重高估模型泛化能力,稳健机器学习评估必须先于临床验证,才能真实评估实际应用价值。

原文摘要 · Abstract (English)

Non-invasive blood glucose level (BGL) estimation from photoplethysmography (PPG) holds great promise for wearable health monitoring, but results across studies are hard to compare due to inconsistent datasets, data leakage, and non-standardized evaluation metrics. We present the first reproducible, extensible evaluation pipeline and use it to reassess five representative PPG-based BGL methods on published datasets under three increasingly strict data-split protocols: random window-level, participant-aware, and leave-some-participants-out (LSPO). Models appeared competitive under random splitting but collapsed under participant-aware and LSPO evaluation, with nearly all yielding near-zero or negative R$^2$ values comparable to a mean-prediction baseline. Critically, across every model and split, over 90% of predictions fell within clinically acceptable zones (Clarke Error Grid A+B), including the baseline. This reveals a fundamental disconnect: clinical zone metrics systematically conceal model failure in this domain. Our findings demonstrate that random train-test splits substantially overestimate the generalization of PPG-based BGL models due to sample-level data leakage, and that robust ML evaluation must precede clinical validation to meaningfully assess real-world utility.

血糖估计PPG模型评估临床验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。