首个针对血糖预测的时序基础模型评测基准,验证了预训练模型在少样本下的强泛化能力。
GlucoFM-Bench: Benchmarking Time-Series Foundation Models for Blood Glucose Forecasting

- 构建跨15个数据集、覆盖1117人的统一评测框架
- 预训练模型在零样本下表现接近全量训练模型(差距<5%)
- 轻量LSTM在有充足数据时仍最优,提示场景适配重要性
血糖预测模型是现代糖尿病管理的核心,可靠的短期预测可推动主动干预与自动胰岛素输送,降低低/高血糖风险。由于糖尿病人群生理差异大,该任务具独特挑战。尽管传统机器学习和深度学习模型已有广泛研究,但近期时序基础模型(TSFMs)在此领域仍缺乏系统评估。为此,我们提出GlucoFM-Bench,一个综合性基准,评估前沿TSFMs与监督式深度学习模型在血糖预测中的表现。涵盖8种代表性架构,包括预训练TSFMs、时序大语言模型及专用深度学习模型,在15个公开糖尿病相关数据集上测试,涉及1117名1型、2型糖尿病、糖尿病前期及健康人群。在零样本、少样本和全量样本条件下,系统评估上下文长度与预测时长的影响。结果显示,预训练模型(尤其是Chronos-2和TimesFM)在零样本与少样本迁移中表现优异,最佳零样本模型性能仅比最优全量监督模型低5%以内。然而,当任务数据充足时,轻量级LSTM仍为最强,性能优于TSFMs达4%-21%。分层分析揭示1型糖尿病群体及低/高血糖区间的持续挑战,强调需超越平均误差指标的评估。GlucoFM-Bench为血糖预测中基础模型的评估、比较与改进提供了标准化、可复现的基础。
原文摘要 · Abstract (English)
Blood glucose forecasting models are foundational for modern diabetes management systems, as reliable short-term predictions can enable proactive interventions, support automated insulin delivery, and reduce the risk of hypo- and hyperglycemic events. From a modeling perspective, glucose forecasting poses unique challenges due to heterogeneous physiological dynamics across diabetes populations. Traditional machine learning and deep learning models have been extensively evaluated for glucose prediction, yet recent time-series foundation models (TSFMs) remain much less studied in this setting. To bridge this gap, we present GlucoFM-Bench, a comprehensive benchmark evaluating state-of-the-art TSFMs alongside supervised deep learning models for blood glucose forecasting. We assess eight representative architectures, including pre-trained TSFMs, time-series large language models, and task-specific deep learning models, across 15 publicly available diabetes-relevant datasets comprising 1,117 individuals with type 1 diabetes, type 2 diabetes, prediabetes, and no diabetes. Models are evaluated under zero-shot, few-shot, and full-shot protocols, with systematic variation in context length and prediction horizon. Across datasets, pre-trained TSFMs, especially Chronos-2 and TimesFM, show strong zero-shot and few-shot transfer, with the best zero-shot model performing within 5% of the best full-shot supervised model. Yet, when task-specific data are abundant, a lightweight LSTM remains strongest, outperforming TSFMs by 4--21% under full-shot training. Stratified analyses reveal persistent challenges in T1D cohorts and hypo-/hyperglycemic ranges, highlighting the need for evaluation beyond aggregate error metrics. Together, GlucoFM-Bench provides a standardized and reproducible foundation for evaluating, comparing, and improving foundation models for blood glucose forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。