arXiv:2609.08772cs.AI2026-09

不同信息表达方式显著影响大模型预测糖尿病血糖事件的效果。

It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction

论文配图:It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction
图 1 · 摘自论文原文
  • 用不同文本形式呈现生理数据,测试大模型在血糖波动预测中的表现。
  • 零样本和少样本下,大模型对低血糖预测优于传统模型,但高血糖预测仍逊色。
  • 信息表达方式比增加额外上下文更关键,适合医疗时序预测研究者参考。

大型语言模型(LLMs)在生理时序预测中的应用日益增多,但其效果不仅取决于模型本身,还受推理时生理信息表示方式的影响。本研究考察了基于提示的通用大模型在1型糖尿病患者餐后高血糖与低血糖预测中的表现。基于OhioT1DM数据集,评估了多种开源大模型在零样本与少样本推理下的性能,预测时长分别为30、60和90分钟。分析对比了从仅葡萄糖数据到包含胰岛素、进餐、碳水化合物及运动等上下文变量的多种文本化信息表示方式。结果表明,模型表现具有任务依赖性:传统监督模型在高血糖预测中表现最优,而最佳提示配置的大模型在所有预测时长下均提升了低血糖预测效果。提示推理的有效性强烈依赖于信息表达方式,而增加上下文信息并未带来系统性改进。总体而言,这些发现强调了在基于提示的大模型血糖事件预测中,生理信息表示是核心设计因素。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only on the model itself, but also on how physiological information is represented and presented at inference time. This study investigates prompt-based general-purpose LLMs for postprandial hyperglycemia and hypoglycemia prediction in individuals with type 1 diabetes. Using the OhioT1DM dataset, we evaluate multiple open-weight LLMs under zero-shot and few-shot inference across prediction horizons of 30, 60, and 90 minutes. The analysis varies both the textual representation of the available physiological information and the amount of information exposed to the model, ranging from glucose observations alone to derived descriptors and additional contextual variables related to insulin, meals, carbohydrates, and physical activity. Performance is compared with conventional patient-specific supervised models and with Gluco-LLM, a language-model-based architecture explicitly adapted to glucose time-series forecasting. Results show a marked task-dependent behavior. Conventional supervised models achieve the strongest performance for hyperglycemia prediction, whereas the best observed prompt-based LLM configurations improve performance for hypoglycemia across all investigated horizons. The effectiveness of prompt-based inference is also strongly influenced by how physiological information is represented, while providing additional contextual information does not lead to a systematic improvement. Overall, these findings highlight physiological information representation as a central design factor in prompt-based LLM approaches to glycemic-event prediction.

大模型血糖预测信息表示医疗时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。