用强化学习为重症患者定制营养方案,降低死亡率并提升代谢稳定。
DeepEN: A Deep Reinforcement Learning Framework for Personalized Enteral Nutrition in Critical Care
- 基于电子病历数据,4小时更新一次个性化营养目标。
- 死亡率降至18.8%,较临床实践降低4个百分点,代谢指标更稳定。
- 推荐依据器官功能等生理指标,非固定用药规则,适合重症监护场景。
重症监护中肠内营养(EN)供给仍缺乏个性化,难以应对动态代谢需求。本文提出DeepEN,一种基于强化学习的个性化营养优化框架,利用MIMIC-IV数据库中超过11,000名患者的电子健康记录数据,生成每4小时一次的个体化热量、蛋白质和液体目标。状态表示包含人口学、共病、生命体征、实验室值及近期干预信息。奖励函数结合生物标志物稳定性与长期生存率,采用带保守Q学习正则化的双深度双网络策略,实现安全离线训练。结果表明,DeepEN估计策略价值最高($V^π=9.48$),校准后死亡率为18.8±1.0%,比临床实践(22.8%)降低4.0个百分点;同时在葡萄糖、磷酸盐、钠等代谢指标达标率上表现最优。偏离DeepEN策略与死亡率及生化不稳显著相关,而偏离随机策略无此关联。可解释性分析显示,推荐依赖于器官功能与代谢状态等生理指标,而非静态剂量规则。结论表明,保守离线强化学习可用于安全、个性化的重症营养优化,推动数据驱动方法与指南相结合。
原文摘要 · Abstract (English)
Objective: Enteral nutrition (EN) delivery in the ICU remains suboptimal due to limited personalization and uncertainty regarding appropriate calorie, protein, and fluid targets under dynamic metabolic demands. We introduce DeepEN, a reinforcement learning (RL) framework for personalized EN optimization using electronic health record data. Methods: DeepEN was trained on over 11,000 ICU patients from MIMIC-IV to generate 4-hourly, patient-specific caloric, protein, and fluid targets. The state representation incorporated demographics, comorbidities, vital signs, laboratory values, and recent interventions. A physiologically aligned reward framework balanced biomarker stability with long-term survival. Policy learning employed a dueling double deep Q-network with Conservative Q-Learning regularization to enable safe offline training. Results: DeepEN achieved the highest estimated policy value ($V^π= 9.48$) and the lowest calibrated mortality (18.8 +/- 1.0%), representing a 4.0 percentage-point absolute reduction compared with clinician practice (22.8%). The policy also demonstrated superior metabolic stability, achieving the highest proportion of glucose, phosphate, and sodium values within target range. Furthermore, deviation from the DeepEN policy was independently associated with increased mortality and biomarker instability, whereas deviation from a random policy showed no such association. Interpretability analyses further indicated that recommendations were conditioned on physiologically relevant markers of organ function and metabolic status rather than static dosing heuristics. Conclusion: DeepEN demonstrates the feasibility of conservative offline RL for safe, individualized EN optimization, highlighting the potential of data-driven personalization to complement guideline-based approaches in critical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。