arXiv:2608.16482cs.AIcs.LG2026-08

用强化学习优化脓毒症患者液体和升压药用药,提升治疗决策可靠性。

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

论文配图:Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
图 1 · 摘自论文原文
  • 构建基于状态-动作网格的马尔可夫决策模型,用策略迭代求解最优用药策略。
  • 两种离线评估方法均显示新策略优于医生实践(得分50.8与46.8,医生为38.2)。
  • 结果稳定且贴近临床实际,适合用于辅助重症监护决策支持系统。

脓毒症患者的静脉输液和升压药给药是受临床判断主导的序列决策过程,适合作为从历史医疗数据中学习强化学习策略的场景。由于无法在真实患者上测试所学策略,必须采用离线策略评估,但此类评估易产生偏差和乐观估计。本研究通过结合离线策略估计、可靠性诊断与医生一致性分析,建立透明验证框架。基于MIMIC-IV数据库中36,872例脓毒症重症监护住院记录,将用药建模为包含1,000个状态与25种动作的离散马尔可夫决策过程,采用五乘五的液体与升压药水平网格定义状态空间,并通过策略迭代求解。使用随机森林估计临床行为策略,显著缓解了有效样本量崩溃问题(ESS从4.0提升至50.1)。采用加权重要性采样(WIS)与拟合Q评价(FQE)两种估计器评估策略,以ESS和医生共识作为可靠性检验。变量选择发现状态构成比大小更重要。两种估计均显示所学策略优于医生表现(WIS 50.8,FQE 46.8 vs 医生38.2,ESS 50.1),且与现有实践差异较小(总变差0.18),更倾向于减少液体输入。这些回顾性单中心离线评估结果支持该策略作为临床实践的合理优化,并推动其作为分歧驱动型临床决策支持工具的进一步评估。

原文摘要 · Abstract (English)

The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.

强化学习重症监护脓毒症离线评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。