构建真实重症患者胰岛素管理数据集,支持离线强化学习研究
Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning

- 基于MIMIC-IV构建含37.5万条真实医嘱的不规则时间数据集
- 在12,209名患者上验证模型性能,突破传统固定时间步限制
- 为重症监护胰岛素管理提供可复现的评估基准,适合临床AI研究者
离线强化学习(ORL)有望利用历史电子健康记录(EHR)数据提升临床决策质量。当前该领域的训练与评估普遍依赖将时间序列数据离散化为固定、规律的时间间隔,这会生成虚构的临床场景,损害回顾性模型评估的泛化能力。本文提出Insulin4RL,一个来自MIMIC-IV的真实临床轨迹数据集,具有自然不规则的输入与动作采样。该数据集包含12,209名需胰岛素输注调整的重症患者,共375,000余条标注决策。它可用于研究在真实临床采样假设下的ORL模型表现。我们描述了数据集结构与特性,提供了基于无模型离线强化学习的基线性能指标,并建立标准化的拟合Q评估协议。最后,提出了未来可利用该资源探索的研究方向。
原文摘要 · Abstract (English)
Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data. Current training and evaluative practices in this field rely heavily on EHR datasets that have been temporally discretised into fixed, regular time intervals. Discretisation creates fictional representations of complex clinical scenarios and compromises the generalisability of retrospective model evaluations. In this paper, we introduce Insulin4RL, a healthcare ORL dataset featuring naturally irregular inputs and actions from real clinical trajectories. Derived from MIMIC-IV, Insulin4RL comprises over 375,000 labelled decisions across 12,209 patients requiring insulin infusion titration in the Intensive Care Unit. The dataset can thus be used for research into ORL model performance under realistic clinical sampling assumptions. We provide a description of the dataset's structure and characteristics, baseline performance metrics using model-free offline reinforcement learning, and a standardised evaluation protocol using fitted Q-evaluation. We conclude with suggested areas for future research that could be addressed using this resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。