arXiv:2409.13299cs.LGcs.AI2024-09被引 1

用离线强化学习模拟医生用药意图,优化肝素剂量决策。

OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment

  • 从有限临床数据中学习参数化奖励函数,捕捉医生治疗意图。
  • 在肝素治疗任务中,新策略显著提升aPTT指标达标率。
  • 方法可推广至其他药物的强化学习给药决策,实用性强。

精准给药在整体治疗过程中至关重要。现有研究多基于强化学习制定最优给药策略,但仅依赖少数显式奖励函数难以覆盖患者个体差异,且临床用药繁多,难以为每种药物定制奖励函数。为此,本文提出离线模型引导的奖励学习方法(OMG-RL),通过离线逆强化学习,从有限数据中学习能反映专家意图的参数化奖励函数,从而优化智能体策略。我们在肝素给药任务上验证该方法,结果表明,OMG-RL策略不仅在学习到的奖励网络上表现优异,且在关键监测指标活化部分凝血酶原时间(aPTT)方面也获得显著提升,说明其充分体现了临床医生的治疗意图。该方法可广泛应用于肝素给药及其他基于强化学习的药物剂量决策任务。

原文摘要 · Abstract (English)

Accurate medication dosing holds an important position in the overall patient therapeutic process. Therefore, much research has been conducted to develop optimal administration strategy based on Reinforcement learning (RL). However, Relying solely on a few explicitly defined reward functions makes it difficult to learn a treatment strategy that encompasses the diverse characteristics of various patients. Moreover, the multitude of drugs utilized in clinical practice makes it infeasible to construct a dedicated reward function for each medication. Here, we tried to develop a reward network that captures clinicians' therapeutic intentions, departing from explicit rewards, and to derive an optimal heparin dosing policy. In this study, we introduce Offline Model-based Guided Reward Learning (OMG-RL), which performs offline inverse RL (IRL). Through OMG-RL, we learn a parameterized reward function that captures the expert's intentions from limited data, thereby enhancing the agent's policy. We validate the proposed approach on the heparin dosing task. We show that OMG-RL policy is positively reinforced not only in terms of the learned reward network but also in activated partial thromboplastin time (aPTT), a key indicator for monitoring the effects of heparin. This means that the OMG-RL policy adequately reflects clinician's intentions. This approach can be widely utilized not only for the heparin dosing problem but also for RL-based medication dosing tasks in general.

强化学习医疗决策肝素给药逆强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。