arXiv:2605.09818cs.LG2026-05

用强化学习压缩慢性病控制时间,提升治疗效率。

Learning to Compress Time-to-Control: A Reinforcement Learning Framework for Chronic Disease Management

  • 设计双环架构,结合偏好学习与强化学习,优化慢性病管理策略。
  • 在糖尿病模拟中,性能比基线高15个百分点,显著缩短控制时间。
  • 适合医疗决策系统研发者及临床研究者参考。

强化学习在医疗领域应用效果参差不齐,主要受限于奖励稀疏、离线评估不可靠以及部署与仿真之间的差距。我们认为,慢性病管理在结构上比急性病症更适合作为强化学习的应用场景,但前提是问题需针对慢性病特性进行形式化建模。本文提出一种新形式化框架:代理目标是压缩‘控制时间’(TTC),采用符合CMS ACCESS模型的分层奖励机制。两个来自前期偏好学习研究的关键参数——执行强度ε(限制动作可用性)和临床能力κ(加权离线数据转移)——作为核心结构要素,将偏好学习与强化学习耦合为双环架构。在高血压和2型糖尿病的合成状态机模拟中,基于能力加权的离线强化学习比均匀加权方法和行为策略在T2D TTC上提升15个百分点;而标准均匀加权方法甚至不如异质行为策略。ε感知策略具有跨部署环境泛化能力,而ε无感知策略则不具备。

原文摘要 · Abstract (English)

Reinforcement learning (RL) in healthcare has had mixed results, with reward sparsity, unreliable off-policy evaluation, and deployment-simulation gap as recurring failure modes. We argue that chronic disease management is structurally a more tractable RL setting than the acute-care problems the field has primarily studied, but only if the problem is formalized to exploit chronic care's properties. We propose such a formalization. The agent's objective is to compress time-to-control (TTC) under a tiered reward calibrated to the CMS ACCESS Model. Two quantities from our companion preference-learning paper [Singh et al. 2026] enter as load-bearing structural elements: the execution intensity εbounds action availability under a constrained Markov Decision Process, and the clinician capability κweights offline-data transitions during RL training. Together they couple preference learning and RL into a two-loop architecture. We present simulation results on synthetic state machines for hypertension and type 2 diabetes. Capability-weighted offline RL outperforms uniform-weighted offline RL and the behavior policy by 15 percentage points on T2D TTC; the uniform-weighted formulation (the standard in existing healthcare RL) underperforms even the heterogeneous behavior policy. \Epsilon-aware policies generalize across deployment regimes while ε-naive policies do not.

强化学习慢性病管理医疗决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。