arXiv:2604.00385cs.LG2026-04

用强化学习为糖尿病患者提供个性化行为建议,提升血糖控制效果。

GUIDE: Reinforcement Learning for Behavioral Action Support in Type 1 Diabetes

  • 设计结构化行为动作空间,结合胰岛素与饮食干预
  • 在25名患者中实现85.49%的血糖达标时间,低血糖风险可控
  • 策略保留患者原有行为模式,适合临床个性化管理

1型糖尿病管理需持续调整胰岛素和生活方式以维持血糖在安全范围。尽管自动化胰岛素输送(AID)系统改善了血糖控制,许多患者仍无法达到临床目标,亟需新方法提升血糖管理效果。现有强化学习(RL)方法多仅关注胰岛素调节,缺乏行为建议。为此,我们提出GUIDE,一种基于强化学习的决策支持框架,可补充AID技术,提供预防高/低血糖事件的行为建议。GUIDE生成包含干预类型、强度和时机的结构化动作,涵盖胰岛素注射与碳水化合物摄入。该框架集成基于真实连续血糖监测数据训练的个体化血糖预测器,并支持离线与在线强化学习算法。我们在25名1型糖尿病患者中评估了多种方法,发现CQL-BC算法平均血糖达标时间达85.49%,同时保持低低血糖暴露。行为相似性分析显示,该策略在患者间保持了0.87±0.09的平均余弦相似度,表明其行为模式与真实患者习惯高度一致。结果表明,采用结构化行为动作空间的保守离线强化学习可为个性化糖尿病管理提供临床有效且行为合理的决策支持。

原文摘要 · Abstract (English)

Type 1 Diabetes (T1D) management requires continuous adjustment of insulin and lifestyle behaviors to maintain blood glucose within a safe target range. Although automated insulin delivery (AID) systems have improved glycemic outcomes, many patients still fail to achieve recommended clinical targets, warranting new approaches to improve glucose control in patients with T1D. While reinforcement learning (RL) has been utilized as a promising approach, current RL-based methods focus primarily on insulin-only treatment and do not provide behavioral recommendations for glucose control. To address this gap, we propose GUIDE, an RL-based decision-support framework designed to complement AID technologies by providing behavioral recommendations to prevent abnormal glucose events. GUIDE generates structured actions defined by intervention type, magnitude, and timing, including bolus insulin administration and carbohydrate intake events. GUIDE integrates a patient-specific glucose level predictor trained on real-world continuous glucose monitoring data and supports both offline and online RL algorithms within a unified environment. We evaluate both off-policy and on-policy methods across 25 individuals with T1D using standardized glycemic metrics. Among the evaluated approaches, the CQL-BC algorithm demonstrates the highest average time-in-range, reaching 85.49% while maintaining low hypoglycemia exposures. Behavioral similarity analysis further indicates that the learned CQL-BC policy preserves key structural characteristics of patient action patterns, achieving a mean cosine similarity of 0.87 $\pm$ 0.09 across subjects. These findings suggest that conservative offline RL with a structured behavioral action space can provide clinically meaningful and behaviorally plausible decision support for personalized diabetes management.

糖尿病管理强化学习行为推荐个性化医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。