让糖尿病患者用偏好数据训练智能胰岛素方案,更安全更个性。
Flexible Blood Glucose Control: Offline Reinforcement Learning from Human Feedback
- 用患者标注的血糖目标生成奖励信号,指导强化学习。
- 比商用系统降低15%血糖风险,餐后血糖达标时间提升10%。
- 支持患者快速调整用药策略,适合临床个性化管理场景。
强化学习在模拟1型糖尿病(T1D)患者中已成功实现胰岛素剂量自动化,但难以融入患者经验和偏好。本文提出PAINT(T1D胰岛素控制偏好适配),一种基于患者记录的离线强化学习框架。PAINT采用基于草图的奖励学习方法,将历史数据标注为连续奖励信号以反映患者的期望结果。标注数据训练奖励模型,指导一种新型安全约束离线强化学习算法,该算法限制动作范围在安全策略内,并通过滑动标尺实现偏好调节。仿真评估显示,仅需对期望状态进行简单标注,PAINT即可达成常见血糖目标,较商业基准降低15%糖化风险。动作标注还可融合患者经验,实现餐前预判(餐后血糖达标时间提升10%)及应对设备错误(错误后方差降低1.6%)。上述结果在有限样本、标注误差和患者间差异等真实条件下依然成立。本研究展示了PAINT在真实T1D管理中的潜力,也适用于任何需快速精准学习偏好且受安全约束的任务。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated success in automating insulin dosing in simulated type 1 diabetes (T1D) patients but is currently unable to incorporate patient expertise and preference. This work introduces PAINT (Preference Adaptation for INsulin control in T1D), an original RL framework for learning flexible insulin dosing policies from patient records. PAINT employs a sketch-based approach for reward learning, where past data is annotated with a continuous reward signal to reflect patient's desired outcomes. Labelled data trains a reward model, informing the actions of a novel safety-constrained offline RL algorithm, designed to restrict actions to a safe strategy and enable preference tuning via a sliding scale. In-silico evaluation shows PAINT achieves common glucose goals through simple labelling of desired states, reducing glycaemic risk by 15% over a commercial benchmark. Action labelling can also be used to incorporate patient expertise, demonstrating an ability to pre-empt meals (+10% time-in-range post-meal) and address certain device errors (-1.6% variance post-error) with patient guidance. These results hold under realistic conditions, including limited samples, labelling errors, and intra-patient variability. This work illustrates PAINT's potential in real-world T1D management and more broadly any tasks requiring rapid and precise preference learning under safety constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。