用稳定编码器提升脓毒症强化学习,让模型更懂病人病情严重程度。
Stable CDE Autoencoders with Acuity Regularization for Offline Reinforcement Learning in Sepsis Treatment
- 用控制微分方程建模不规则医疗时间序列,保证训练稳定
- 结合临床评分相关性正则化,使表示与病情严重度强关联
- 在MIMIC-III数据上实现超过0.9的加权重要性采样回报
脓毒症治疗的有效强化学习依赖于从不规则ICU时间序列中学习稳定且具有临床意义的状态表示。以往工作虽探索了表示学习,但忽视了序列表示训练不稳定对策略性能的负面影响。本研究发现,当满足两个关键条件时,控制微分方程(CDE)状态表示可生成优异的强化学习策略:(1) 通过早停或稳定化方法确保训练稳定;(2) 通过与临床评分(SOFA、SAPS-II、OASIS)的相关性正则化,强制生成病情严重度感知表示。在MIMIC-III脓毒症队列上的实验表明,稳定CDE自编码器生成的表示与病情严重度高度相关,并使强化学习策略取得优越性能(加权重要性采样回报 > 0.9)。相比之下,不稳定CDE表示导致表示退化和策略失败(回报 ~ 0)。潜空间可视化显示,稳定CDE不仅能分离生存与非生存轨迹,还呈现清晰的病情严重度梯度,而不稳定训练无法捕捉这些模式。这些发现为使用CDE编码不规则医疗时间序列提供了实用指导,强调了序列表示学习中训练稳定性的重要性。
原文摘要 · Abstract (English)
Effective reinforcement learning (RL) for sepsis treatment depends on learning stable, clinically meaningful state representations from irregular ICU time series. While previous works have explored representation learning for this task, the critical challenge of training instability in sequential representations and its detrimental impact on policy performance has been overlooked. This work demonstrates that Controlled Differential Equations (CDE) state representation can achieve strong RL policies when two key factors are met: (1) ensuring training stability through early stopping or stabilization methods, and (2) enforcing acuity-aware representations by correlation regularization with clinical scores (SOFA, SAPS-II, OASIS). Experiments on the MIMIC-III sepsis cohort reveal that stable CDE autoencoder produces representations strongly correlated with acuity scores and enables RL policies with superior performance (WIS return $> 0.9$). In contrast, unstable CDE representation leads to degraded representations and policy failure (WIS return $\sim$ 0). Visualizations of the latent space show that stable CDEs not only separate survivor and non-survivor trajectories but also reveal clear acuity score gradients, whereas unstable training fails to capture either pattern. These findings highlight practical guidelines for using CDEs to encode irregular medical time series in clinical RL, emphasizing the need for training stability in sequential representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。