arXiv:2606.01051cs.LG2026-06

让医疗治疗决策自动优化时间与强度,兼顾连续安全与疗效。

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

论文配图:Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment
图 1 · 摘自论文原文
  • 将治疗看作可变时长的策略选项,动态决定何时干预、如何用药。
  • 在连续时间中保证安全,实验显示比固定间隔方案更安全有效。
  • 适合需精准控制治疗节奏的慢性病或重症管理场景。

动态医疗治疗需同时决定治疗强度和干预时机,而患者状态持续演变,不良事件可能发生在临床交互之间。现有方法多假设固定干预周期,或仅在离散时刻保障安全。本文提出交互受限的安全连续时间强化学习框架,联合优化治疗实施与临床交互时机,满足轨迹级安全约束。核心思想是将连续时间治疗问题重构为基于选项的半马尔可夫决策过程,每个选项定义一个持续时间内的治疗策略。我们设计了一种安全强化机制,证明在交互时刻施加适当约束即可以高概率保障全连续轨迹的安全性。进一步建立了从记录治疗轨迹中学习策略的有限样本保证,并提出一种实用的数据驱动保守代理函数。实验表明,所提出的自适应交互时机机制在多种安全策略优化方法下均优于等距交互方案,在安全性和治疗效果上均有提升。

原文摘要 · Abstract (English)

Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical interactions. Most existing treatment learning methods assume fixed schedules or enforce safety only at discrete decision points. We propose Interaction-Limited Safe Continuous-Time Reinforcement Learning, a framework that jointly optimizes treatment administration and clinical interaction timing under trajectory-level safety constraints. Our key idea is to reformulate the continuous time treatment problem as an option-based semi-Markov decision process, where each option specifies a continuous-time treatment policy and its duration. We develop a safety-tightening mechanism showing that suitably constructed constraints at interaction times guarantee safety over the full continuous-time trajectory with high probability. We further establish finite-sample guarantees for policy learning from logged treatment trajectories and introduce a practical data-driven conservative surrogate. Experiments show that the proposed adaptive interaction-timing mechanism improves both safety and treatment effectiveness over equidistant interaction schemes across different safe policy optimization methods.

医疗决策连续时间强化学习安全约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。