arXiv:2511.20913cs.LGcs.AI2025-11被引 4

对比不同时间步长对脓毒症强化学习治疗效果的影响

Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment

  • 测试1、2、4、8小时四种时间步长,统一训练流程确保公平比较
  • 1小时和2小时步长在行为克隆与策略训练中表现最佳且更稳定
  • 揭示时间步长是医疗强化学习的核心设计选择,挑战传统4小时设定

现有脓毒症管理的强化学习研究多采用4小时为时间步长的数据聚合方式。尽管有人质疑这种粗粒度可能扭曲患者动态并导致次优治疗策略,但其实际影响尚不明确。本文通过在相同离线强化学习流程下,对Δt=1、2、4、8小时四种时间步长进行实证对比,设计动作重映射方法以实现跨步长评估,并在两种策略学习设置下开展跨Δt模型选择。研究旨在量化时间步长对状态表示学习、行为克隆、策略训练及离线评估的影响。结果表明,性能趋势随学习设置而异,使用静态行为策略时,1小时和2小时步长的策略整体表现最优且最稳定。本工作强调时间步长是医疗离线强化学习的关键设计因素,为突破传统4小时设定提供实证支持。

原文摘要 · Abstract (English)

Existing studies on reinforcement learning (RL) for sepsis management have mostly followed an established problem setup, in which patient data are aggregated into 4-hour time steps. Although concerns have been raised regarding the coarseness of this time-step size, which might distort patient dynamics and lead to suboptimal treatment policies, the extent to which this is a problem in practice remains unexplored. In this work, we conducted empirical experiments for a controlled comparison of four time-step sizes ($Δt\!=\!1,2,4,8$ h) on this domain, following an identical offline RL pipeline. To enable a fair comparison across time-step sizes, we designed action re-mapping methods that allow for evaluation of policies on datasets with different time-step sizes, and conducted cross-$Δt$ model selections under two policy learning setups. Our goal was to quantify how time-step size influences state representation learning, behavior cloning, policy training, and off-policy evaluation. Our results show that performance trends across $Δt$ vary as learning setups change, while policies learned at finer time-step sizes ($Δt = 1$ h and $2$ h) using a static behavior policy achieve the overall best performance and stability. Our work highlights time-step size as a core design choice in offline RL for healthcare and provides evidence supporting alternatives beyond the conventional 4-hour setup.

强化学习医疗决策时间步长脓毒症

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。