arXiv:2410.07525cs.LGcs.AI2024-10被引 10

通过历史数据学习医疗决策中的安全约束,避免药物过量等危险行为。

Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare

  • 用因果注意力机制融合历史治疗记录,建模非马尔可夫约束。
  • 生成式世界模型增强数据,模拟危险决策路径并降低出错概率。
  • 适合医疗决策系统开发者,尤其关注安全性的离线强化学习应用。

将强化学习应用于医疗领域时,常因忽视常识性约束而产生不安全的治疗决策,如药物过量或突变剂量。约束强化学习(CRL)是保障安全的自然选择,但医疗场景中精确设定代价函数极为困难。逆向约束强化学习(ICRL)通过专家示范推断约束,但传统方法依赖交互环境中的马尔可夫决策,与实际医疗中基于离线历史数据做决策的需求不符。为此,我们提出约束变换器(Constraint Transformer, CT):1)采用因果注意力机制融合历史决策与观测,结合非马尔可夫层加权约束以捕捉关键状态;2)引入生成式世界模型进行探索性数据增强,使离线强化学习能模拟不安全决策序列。在多个医疗场景中,实验证明CT能有效识别不安全状态,并生成逼近更低死亡率的策略,显著降低不安全行为的发生概率。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) applied in healthcare can lead to unsafe medical decisions and treatment, such as excessive dosages or abrupt changes, often due to agents overlooking common-sense constraints. Consequently, Constrained Reinforcement Learning (CRL) is a natural choice for safe decisions. However, specifying the exact cost function is inherently difficult in healthcare. Recent Inverse Constrained Reinforcement Learning (ICRL) is a promising approach that infers constraints from expert demonstrations. ICRL algorithms model Markovian decisions in an interactive environment. These settings do not align with the practical requirement of a decision-making system in healthcare, where decisions rely on historical treatment recorded in an offline dataset. To tackle these issues, we propose the Constraint Transformer (CT). Specifically, 1) we utilize a causal attention mechanism to incorporate historical decisions and observations into the constraint modeling, while employing a Non-Markovian layer for weighted constraints to capture critical states. 2) A generative world model is used to perform exploratory data augmentation, enabling offline RL methods to simulate unsafe decision sequences. In multiple medical scenarios, empirical results demonstrate that CT can capture unsafe states and achieve strategies that approximate lower mortality rates, reducing the occurrence probability of unsafe behaviors.

医疗AI约束强化学习离线决策安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。