用时序差分学习提升重症监护死亡预测的稳定性。
Robust Real-Time Mortality Prediction in the Intensive Care Unit using Temporal Difference Learning
- 将时序差分学习引入不规则采样医疗数据,构建半马尔可夫奖励模型。
- 在多个外部数据集上验证,模型鲁棒性显著优于传统监督学习。
- 适合高波动、不规则时间序列的临床预后预测场景。
利用监督学习预测长期患者结局是一项挑战,主要源于个体轨迹的高方差,易导致模型过拟合。时序差分(TD)学习作为强化学习常见技术,可通过泛化状态转移模式而非终端结果来降低方差。然而,其在医疗领域的应用受限于对患者状态的强假设,且缺乏与传统监督学习方法在长期健康结局预测中的对比研究。本研究提出一种基于半马尔可夫奖励过程的框架,用于在实时不规则采样时间序列数据中应用TD学习。在重症监护死亡预测任务中评估该框架,结果表明,在多种外部数据集上,该方法相比标准监督学习具备更强的鲁棒性,且性能保持稳定。该方法为处理高方差、不规则时间序列数据的患者结局预测提供了更可靠的新路径。
原文摘要 · Abstract (English)
The task of predicting long-term patient outcomes using supervised machine learning is a challenging one, in part because of the high variance of each patient's trajectory, which can result in the model over-fitting to the training data. Temporal difference (TD) learning, a common reinforcement learning technique, may reduce variance by generalising learning to the pattern of state transitions rather than terminal outcomes. However, in healthcare this method requires several strong assumptions about patient states, and there appears to be limited literature evaluating the performance of TD learning against traditional supervised learning methods for long-term health outcome prediction tasks. In this study, we define a framework for applying TD learning to real-time irregularly sampled time series data using a Semi-Markov Reward Process. We evaluate the model framework in predicting intensive care mortality and show that TD learning under this framework can result in improved model robustness compared to standard supervised learning methods. and that this robustness is maintained even when validated on external datasets. This approach may offer a more reliable method when learning to predict patient outcomes using high-variance irregular time series data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。