arXiv:2510.01508cs.LG2025-10

用强化学习设计可执行的双血管活性药剂量方案,提升重症监护可操作性。

Realistic CDSS Drug Dosing with End-to-end Recurrent Q-learning for Dual Vasopressor Control

  • 设计混合离散、连续与方向性给药空间,确保策略可执行
  • 在eICU和MIMIC数据集上实现3倍以上奖励提升
  • 适合临床部署,结果符合现有医疗指南

强化学习在临床决策支持系统中的应用常因推荐不可执行的给药方案而受质疑。本文提出一种端到端离线强化学习框架,用于重症监护室双血管活性药物管理,通过合理的动作空间设计直接解决该问题。方法融合离散、连续及方向性给药策略,结合保守Q-learning,并引入新型循环建模机制,利用回放缓冲区捕捉危重患者时序数据中的动态依赖关系。在不同动作空间设置下对去甲肾上腺素给药策略的对比分析表明,所设计的动作空间显著提升策略可解释性,促进临床采纳,同时保持疗效。在eICU和MIMIC数据集上的实证结果表明,动作空间设计深刻影响学习到的行为策略。相比基线方法,本方法实现超过3倍的预期奖励提升,且与现有临床指南高度一致。

原文摘要 · Abstract (English)

Reinforcement learning (RL) applications in Clinical Decision Support Systems (CDSS) frequently encounter skepticism because models may recommend inoperable dosing decisions. We propose an end-to-end offline RL framework for dual vasopressor administration in Intensive Care Units (ICUs) that directly addresses this challenge through principled action space design. Our method integrates discrete, continuous, and directional dosing strategies with conservative Q-learning and incorporates a novel recurrent modeling using a replay buffer to capture temporal dependencies in ICU time-series data. Our comparative analysis of norepinephrine dosing strategies across different action space formulations reveals that the designed action spaces improve interpretability and facilitate clinical adoption while preserving efficacy. Empirical results on eICU and MIMIC demonstrate that action space design profoundly influences learned behavioral policies. Compared with baselines, the proposed methods achieve more than 3x expected reward improvements, while aligning with established clinical protocols.

强化学习临床决策重症监护药物剂量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。