用强化学习优化重症镇痛镇静,兼顾疼痛缓解与出院后30天死亡率。
On Safer Reinforcement Learning for Sedation and Analgesia in Intensive Care
- 基于回顾性数据训练深度强化学习模型,每小时推荐四种药物剂量。
- 联合降低疼痛与死亡率的策略显著提升临床认可度,且与死亡率负相关。
- 考虑长期生存结果能避免高共病患者治疗风险,适合临床决策支持系统使用。
重症监护中疼痛管理常面临权衡困境:治疗不足或过度均可能危及患者安全。以往针对镇痛镇静的强化学习研究未考虑患者生存率或部分可观测性问题。为此,我们构建了一个离线深度强化学习框架,基于循环状态表示每小时建议药物剂量。利用MIMIC-IV数据库中47,144例患者的回顾数据,训练了行为正则化的演员-评论家模型,分别以减轻疼痛或同时减轻疼痛与30天出院后死亡率为目标,输出阿片类、丙泊酚、苯二氮䓬类和右美托咪定的连续剂量。尽管两种策略均降低疼痛,但仅关注疼痛的策略与临床医生同意度呈正相关(ρ=0.119,p<0.0001),而联合目标策略则呈负相关(ρ=-0.316,p<0.0001)。差异源于对高共病水平患者的响应不同,表明即使短期目标为主,纳入长期结局对学习更安全的治疗策略至关重要。
原文摘要 · Abstract (English)
Pain management in intensive care usually involves complex trade-offs, since both inadequate and excessive treatment can compromise patient safety. Prior work on reinforcement learning for sedation and analgesia has explored how to optimize these interventions, but has not considered patient survival or partial observability. To investigate the risks of these design choices, we developed an offline deep reinforcement learning framework that suggests hourly medication doses based on recurrent state representations. Using retrospective data from 47,144 ICU stays in the MIMIC-IV database, we trained and evaluated behavior-regularized actor-critic models that prescribe continuous doses of opioids, propofol, benzodiazepines, and dexmedetomidine according to two goals: reduce pain or jointly reduce pain and 30-day post-discharge mortality. Although the two resulting policies were associated with lower pain, clinician agreement with the pain-only policy was positively correlated with mortality ($ρ$=0.119, p<0.0001), while agreement with the joint policy was negatively correlated ($ρ$=-0.316, p<0.0001). We found that such divergence arose from a different response to high levels of comorbidity. This suggests that valuing post-discharge outcomes could be critical for learning safer treatment policies, even if a short-term goal remains the primary objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。