arXiv:2508.17212cs.AI2025-08被引 1

用数字孪生和强化学习实现可实时调整的临床决策支持。

Reinforcement Learning enhanced Online Adaptive Clinical Decision Support via Digital Twin powered Policy and Treatment Effect optimized Reward

  • 用患者数字孪生构建环境,强化学习生成治疗策略。
  • 专家干预率低,系统在保持安全的前提下提升治疗效果。
  • 适合需要快速适应且注重安全的临床场景使用。

临床决策支持需在安全约束下在线自适应。本文提出一种在线自适应工具:强化学习生成策略,患者数字孪生提供仿真环境,治疗效果定义奖励函数。系统从回顾性数据初始化批量约束策略,随后运行流式循环,选择动作、验证安全,并仅在不确定性高时查询专家。不确定性通过五个Q网络组成的紧凑集成,结合动作值方差系数与tanh压缩计算。数字孪生采用有界残差规则更新患者状态。结果模型估计即时临床效果,奖励为相对于保守基准的治疗效果,基于训练集固定z-score归一化。在线更新基于近期数据,采用短周期与指数移动平均。规则式安全门控在动作执行前强制生命体征范围与禁忌症检查。合成临床模拟器实验显示,系统具备低延迟、稳定吞吐量、固定安全下的低专家查询率,并优于标准值基基线的回报。该设计将离线策略转化为持续、医师监督的系统,具备明确控制与快速适应能力。

原文摘要 · Abstract (English)

Clinical decision support must adapt online under safety constraints. We present an online adaptive tool where reinforcement learning provides the policy, a patient digital twin provides the environment, and treatment effect defines the reward. The system initializes a batch-constrained policy from retrospective data and then runs a streaming loop that selects actions, checks safety, and queries experts only when uncertainty is high. Uncertainty comes from a compact ensemble of five Q-networks via the coefficient of variation of action values with a $\tanh$ compression. The digital twin updates the patient state with a bounded residual rule. The outcome model estimates immediate clinical effect, and the reward is the treatment effect relative to a conservative reference with a fixed z-score normalization from the training split. Online updates operate on recent data with short runs and exponential moving averages. A rule-based safety gate enforces vital ranges and contraindications before any action is applied. Experiments in a synthetic clinical simulator show low latency, stable throughput, a low expert query rate at fixed safety, and improved return against standard value-based baselines. The design turns an offline policy into a continuous, clinician-supervised system with clear controls and fast adaptation.

临床决策强化学习数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。