用安全约束的强化学习实现心脏辅助设备自动撤机,提升临床决策可靠性。
Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices
- 基于临床知识设计奖励机制与分布外抑制算法,确保策略安全
- 在真实与合成数据上比基线提升28%奖励与82.6%临床评分
- 适合高风险医疗场景,需结合医学先验与安全约束的研究者
本文研究心源性休克患者机械循环支持(MCS)设备自动化撤机中的序列决策问题。MCS为经皮微轴流泵,可减轻左心室负荷并维持血流,但当前撤机策略因医护团队差异大且缺乏数据驱动方法。传统离线强化学习在该场景面临挑战:禁止在线患者交互、循环系统动态高度不确定(因联合治疗)、数据有限。本文提出端到端机器学习框架,核心贡献为:(1) 临床感知的分布外正则化模型基策略优化(CORMPO),一种密度正则化的离线强化学习算法,结合临床指导的奖励设计;(2) 基于Transformer的概率数字孪生模型,用于策略评估,包含丰富的生理与临床指标。理论证明CORMPO在弱假设下具备性能保证。在真实与合成数据集上,CORMPO相比离线强化学习基线,奖励提升28%,临床指标得分提升82.6%。本方法为高风险医疗应用中安全离线策略学习提供了系统性框架。
原文摘要 · Abstract (English)
We study the sequential decision-making problem for automated weaning of mechanical circulatory support (MCS) devices in cardiogenic shock patients. MCS devices are percutaneous micro-axial flow pumps that provide left ventricular unloading and forward blood flow, but current weaning strategies vary significantly across care teams and lack data-driven approaches. Offline reinforcement learning (RL) has proven to be successful in sequential decision-making tasks, but our setting presents challenges for training and evaluating traditional offline RL methods: prohibition of online patient interaction, highly uncertain circulatory dynamics due to concurrent treatments, and limited data availability. We developed an end-to-end machine learning framework with two key contributions (1) Clinically-aware OOD-regularized Model-based Policy Optimization (CORMPO), a density-regularized offline RL algorithm for out-of-distribution suppression that also incorporates clinically-informed reward shaping and (2) a Transformer-based probabilistic digital twin that models MCS circulatory dynamics for policy evaluation with rich physiological and clinical metrics. We prove that \textsf{CORMPO} achieves theoretical performance guarantees under mild assumptions. CORMPO attains a higher reward than the offline RL baselines by 28% and higher scores in clinical metrics by 82.6% on real and synthetic datasets. Our approach offers a principled framework for safe offline policy learning in high-stakes medical applications where domain expertise and safety constraints are essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。