arXiv:2605.00645cs.LG2026-05

提出评估血糖预测模型真实临床价值的新框架,区分预警与胰岛素决策场景。

From Prediction to Practice: A Task-Aware Evaluation Framework for Blood Glucose Forecasting

  • 按临床任务分设预警和干预两个评估维度,避免传统误差指标误导。
  • 真实数据中部分模型在高危时段召回率仅0.72,误报达每患者日3.1次。
  • 在模拟干预测试中,多数模型无法准确预测胰岛素调整的血糖反应方向。

临床时间序列预测日益用于辅助决策,但传统综合指标可能掩盖模型在关键任务中的实际效用。在高风险场景中,平均误差低的模型仍可能在最需要预警的时段出现严重漏报。本文构建面向血糖预测的任务感知评估框架,聚焦两大下游应用:低血糖早期预警与胰岛素剂量决策支持。针对预警任务,基于三个临床队列的真实数据,采用事件级召回率与每患者日误报数作为指标,反映实际报警负担而非整体准确率。结果显示,尽管某些模型在全集上召回率超0.9,但在餐后胰岛素负荷高峰期,其召回率下降至0.72,且误报达每患者日3.1次,临床后果严重。标准预测评估未检验模型对治疗行为影响的推理能力,而这是支持胰岛素决策的必要条件。因此,我们引入第二个干预性评估分支,使用FDA认可的UVA/Padova模拟器,评估模型在配对事实/反事实情境下对胰岛素方案变更的血糖响应预测能力。结果表明,表现良好的真实数据模型在干预效应的方向、幅度或排序预测上常失败,导致临床成本导向的评分下选择劣质胰岛素剂量。两部分评估共同揭示了预测精度与任务相关实用性之间的系统性差距。本文发布基准测试、公开队列标准化预处理流程及模拟干预数据集,形成可复现的工具包。

原文摘要 · Abstract (English)

Clinical time-series forecasting is increasingly studied for decision support, yet standard aggregate metrics can obscure whether a model is actually useful for the task it is meant to serve. In safety-critical settings, low average error can coexist with dangerous failures in exactly the high-risk regimes that matter most. We present a task-aware evaluation framework for blood glucose forecasting built around two downstream uses: hypoglycemia early warning and insulin dosing decision support. For early warning, we evaluate on real data from three clinical cohorts using event-level recall and false alarms per patient-day, metrics that reflect operational alarm burden rather than aggregate accuracy. We show that models appearing acceptable overall, with recall above 0.9 on the full test set, can fail badly in the post-bolus slice, where insulin-on-board is elevated and missed warnings carry the greatest clinical consequences. Standard forecasting evaluation, however, does not test whether a model can reason about the effects of actions, a requirement for supporting insulin dosing decisions. We therefore add a second, interventional arm using the FDA-accepted UVA/Padova simulator, where we evaluate whether forecasters can predict glucose responses to altered insulin plans in paired factual/counterfactual scenarios. We show that models that look strong on real-data forecasting often fail to predict the direction, magnitude, or ranking of intervention effects, and choose poor insulin doses when evaluated under a clinically motivated cost. Taken together, the two arms reveal a consistent gap between forecasting accuracy and task-relevant usefulness. We release the benchmark, the standardized preprocessing pipeline for public cohorts, and the simulator-based interventional dataset as a reproducible toolkit.

血糖预测评估框架临床决策模拟器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。