用Transformer和安全强化学习优化呼吸机设置,降低肺损伤风险。
Ensuring Safety in Automated Mechanical Ventilation through Offline Reinforcement Learning and Digital Twin Verification
- 用Transformer捕捉患者生理变化的时序特征,提升决策精度。
- 在120例模拟病例中,较现有方法降低38%肺损伤风险。
- 结合数字孪生在线验证,适合重症监护场景的临床决策支持。
机械通气(MV)是急性呼吸衰竭(ARF)患者的重要生命支持手段,但不当设置可能引发呼吸机相关性肺损伤(VILI),且医护人员负担与患者预后直接相关。为实现个性化与自动化,本文提出基于Transformer的保守Q学习(T-CQL)框架,通过Transformer编码器建模患者动态时序特征,结合不确定性量化进行保守自适应正则化以保障安全,并引入一致性正则化提升鲁棒性。设计包含VILI指标与病情严重度评分的临床导向奖励函数。针对传统静态离线数据评估对环境变化不敏感的问题,采用交互式数字孪生系统在床旁实时验证策略性能。实验表明,T-CQL在120例模拟病例中持续优于现有先进离线强化学习方法,显著降低肺损伤风险并提升通气效果,展现了基于Transformer与保守强化学习结合在危重症决策支持中的潜力。
原文摘要 · Abstract (English)
Mechanical ventilation (MV) is a life-saving intervention for patients with acute respiratory failure (ARF) in the ICU. However, inappropriate ventilator settings could cause ventilator-induced lung injury (VILI). Also, clinicians workload is shown to be directly linked to patient outcomes. Hence, MV should be personalized and automated to improve patient outcomes. Previous attempts to incorporate personalization and automation in MV include traditional supervised learning and offline reinforcement learning (RL) approaches, which often neglect temporal dependencies and rely excessively on mortality-based rewards. As a result, early stage physiological deterioration and the risk of VILI are not adequately captured. To address these limitations, we propose Transformer-based Conservative Q-Learning (T-CQL), a novel offline RL framework that integrates a Transformer encoder for effective temporal modeling of patient dynamics, conservative adaptive regularization based on uncertainty quantification to ensure safety, and consistency regularization for robust decision-making. We build a clinically informed reward function that incorporates indicators of VILI and a score for severity of patients illness. Also, previous work predominantly uses Fitted Q-Evaluation (FQE) for RL policy evaluation on static offline data, which is less responsive to dynamic environmental changes and susceptible to distribution shifts. To overcome these evaluation limitations, interactive digital twins of ARF patients were used for online "at the bedside" evaluation. Our results demonstrate that T-CQL consistently outperforms existing state-of-the-art offline RL methodologies, providing safer and more effective ventilatory adjustments. Our framework demonstrates the potential of Transformer-based models combined with conservative RL strategies as a decision support tool in critical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。