构建连续时间医疗强化学习基准,支持不规则治疗与个性化评估
MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

- 用物理信息神经网络构建连续时间患者演化模型
- 支持离线与在线强化学习,可对比时序依赖治疗效果
- 适合关注医疗决策安全与个体差异的研究者
医疗治疗推荐对强化学习(RL)提出多重挑战:患者生理状态在连续时间中演变,测量与干预时间不规则,且治疗反应因人而异。现有RL框架与仿真环境多基于离散时间马尔可夫决策过程(MDP)或部分可观测马尔可夫决策过程(POMDP),采用固定或预设的决策间隔。这使得难以评估RL方法在处理时间间隔相关疾病进展、个性化治疗反应及连续监测点间安全性方面的能力。为弥补这一空白,我们提出MedGym,一个面向动态治疗推荐的基准环境。MedGym在连续时间框架下建模纵向患者演化,并利用物理信息神经网络从临床数据构建可配置的医学强化学习基准。该基准支持离线与在线强化学习,可直接比较离散时间与连续时间方法在不规则治疗时机与个体化动态下的表现。此外,MedGym支持从个性化、轨迹级安全性以及模型驱动离线学习与在线部署性能差距等临床重要角度进行评估。MedGym旨在为连续时间动态治疗推荐提供标准化、可配置的基准,推动医学强化学习方法更真实、更具信息量的评估。
原文摘要 · Abstract (English)
Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals. Existing RL formulations and simulated environments, however, are based on discrete-time MDP or POMDP abstractions with fixed or pre-specified decision intervals. Thus, it remains difficult to evaluate whether RL methods can handle time-interval-dependent disease progression, personalized treatment response, and safety between consecutive measurement points. To address this gap, we introduce MedGym, a benchmark environment for dynamic treatment recommendation. MedGym models longitudinal patient evolution in a continuous-time framework and constructs a configurable medical RL benchmark from clinical data by using Physics-Informed Neural Networks. The resulting benchmark supports both offline and online RL, and enables direct comparison between discrete-time and continuous-time methods under irregular treatment timing and patient-specific dynamics. Besides, MedGym supports evaluation from clinically important perspectives, including personalization, trajectory-level safety, and the performance gap between model-based offline learning and online deployment. By providing a standardized and configurable benchmark for continuous-time dynamic treatment, MedGym aims to facilitate more realistic and informative evaluation of medical RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。