用神经微分方程+安全强化学习,实现精准又安全的糖尿病胰岛素调控。
Integrating Neural Differential Forecasting with Safe Reinforcement Learning for Blood Glucose Regulation
- 用神经微分方程预测餐后血糖变化,结合贝叶斯采样优化胰岛素剂量。
- 在模拟测试中血糖达标时间达87.9%,低血糖时间低于10%。
- 适合关注安全性与个性化调控的医疗AI研究者和开发者。
1型糖尿病的自动胰岛素输注需在不确定的进餐和生理波动下平衡血糖控制与安全性。尽管强化学习(RL)可实现个性化适应,但现有方法难以同时保证安全,存在餐前过量给药或连续纠错等风险。为此,我们提出TSODE,一种融合汤普森采样强化学习与神经常微分方程(NeuralODE)预报器的安全感知控制器。其中,NeuralODE基于预设胰岛素剂量预测短期血糖轨迹,结合置信校准层量化预测不确定性,拒绝或缩放高风险动作。在经FDA批准的UVa/Padova模拟器(成人队列)中,TSODE实现87.9%的时间在目标范围内,低血糖时间低于10%,优于相关基线。结果表明,将自适应强化学习与校准后的NeuralODE预报结合,可实现可解释、安全且鲁棒的血糖调控。
原文摘要 · Abstract (English)
Automated insulin delivery for Type 1 Diabetes must balance glucose control and safety under uncertain meals and physiological variability. While reinforcement learning (RL) enables adaptive personalization, existing approaches struggle to simultaneously guarantee safety, leaving a gap in achieving both personalized and risk-aware glucose control, such as overdosing before meals or stacking corrections. To bridge this gap, we propose TSODE, a safety-aware controller that integrates Thompson Sampling RL with a Neural Ordinary Differential Equation (NeuralODE) forecaster to address this challenge. Specifically, the NeuralODE predicts short-term glucose trajectories conditioned on proposed insulin doses, while a conformal calibration layer quantifies predictive uncertainty to reject or scale risky actions. In the FDA-approved UVa/Padova simulator (adult cohort), TSODE achieved 87.9% time-in-range with less than 10% time below 70 mg/dL, outperforming relevant baselines. These results demonstrate that integrating adaptive RL with calibrated NeuralODE forecasting enables interpretable, safe, and robust glucose regulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。