arXiv:2506.14375cs.LGcs.AI2025-06AAAI被引 6

用离线强化学习优化呼吸机设置,提升重症患者安全与疗效。

Advancing Safe Mechanical Ventilation Using Offline RL With Hybrid Actions and Clinically Aligned Rewards

  • 设计混合动作空间的离线强化学习框架,直接处理连续与离散参数。
  • 可同时优化6个呼吸机参数,显著超过此前2-3个的限制。
  • 结合临床目标的奖励函数,兼顾生存率与生理指标,适合临床部署。

侵入性机械通气(MV)是重症监护病房(ICU)中治疗严重急症患者的救命手段,患者常依赖其呼吸。优化MV设置可降低死亡率、减少呼吸机相关肺损伤、缩短住院时间并缓解医疗资源压力。然而,因患者个体差异大,设置优化仍复杂且易出错。现有离线强化学习方法难以处理MV设置中连续与离散混合的动作空间。若对连续参数离散化,会导致动作空间指数级膨胀,限制可优化参数数量;而将预测结果转回连续值则可能引发分布偏移,影响安全性和性能。为此,在IntelliLung项目中,我们提出一种约束动作空间并采用分解式动作评价器的方法,使可优化设置从以往的2-3个扩展至6个。我们改进了最先进的离线强化学习算法,使其能直接作用于混合动作空间,避免离散化问题。同时引入基于脱机天数和生理目标的临床导向奖励函数,并通过多目标优化选择奖励,实现对各类临床目标更均衡的考量。系统在与医护人员紧密协作下开发,贴合真实临床需求,具备未来部署潜力。

原文摘要 · Abstract (English)

Invasive mechanical ventilation (MV) is a life-sustaining therapy commonly used in the intensive care unit (ICU) for patients with severe and acute conditions. These patients frequently rely on MV for breathing. Given the high risk of death in such cases, optimal MV settings can reduce mortality, minimize ventilator-induced lung injury, shorten ICU stays, and ease the strain on healthcare resources. However, optimizing MV settings remains a complex and error-prone process due to patient-specific variability. While Offline Reinforcement Learning (RL) shows promise for optimizing MV settings, current methods struggle with the hybrid (continuous and discrete) nature of MV settings. Discretizing continuous settings leads to exponential growth in the action space, which limits the number of optimizable settings. Converting the predictions back to continuous can cause a distribution shift, compromising safety and performance. To address this challenge, in the IntelliLung project, we are developing an AI-based approach where we constrain the action space and employ factored action critics. This approach allows us to scale to six optimizable settings compared to 2-3 in previous studies. We adapt SOTA offline RL algorithms to operate directly on hybrid action spaces, avoiding the pitfalls of discretization. We also introduce a clinically grounded reward function based on ventilator-free days and physiological targets. Using multiobjective optimization for reward selection, we show that this leads to a more equitable consideration of all clinically relevant objectives. Notably, we develop a system in close collaboration with healthcare professionals that is aligned with real-world clinical objectives and designed with future deployment in mind.

强化学习呼吸机优化临床AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。