提出临床反事实审计框架,检测重症治疗强化学习中的有害模仿问题。
Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

- 通过生理扰动测试强化学习模型在真实指南下的反应
- 发现某模型在乳酸升高时反而减少升压药使用,违背治疗指南
- 适合医疗AI安全评估、临床决策系统研发者参考
离线强化学习(Offline RL)在优化重症监护治疗决策方面前景广阔,但标准评估指标均方误差(MSE)和拟合Q值评估(FQE)仅衡量行为模仿,无法识别毒性模仿——即代理在舒适护理过渡期错误复制如撤除治疗等有害行为。基于MIMIC-III数据库,本文提出反事实临床审计(CCA)框架,通过符合生存性脓毒症指南(SSC)的生理扰动压力测试强化学习代理。我们评估了医学决策变换器(MedDT)和历史因果变换器(HCT-RL),后者采用因果动作防护、基于倾向性的重要性加权及保守Q学习。CCA结果显示,MedDT在乳酸水平上升时反而降低升压药剂量,违背复苏指南;而HCT-RL保持生理一致性响应。这些发现揭示统计拟合与临床安全性之间存在系统性偏差,支持反事实审计作为医疗强化学习的必要评估标准。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions. Using the MIMIC-III database, we propose the Counterfactual Clinical Audit (CCA) framework, which stress-tests RL agents through physiological perturbations anchored in Surviving Sepsis Campaign (SSC) guidelines. We audit a Medical Decision Transformer (MedDT) and a Historical Causal Transformer (HCT-RL), the latter employing Causal Action Shielding, propensity-based importance weighting, and Conservative Q-Learning. CCA reveals that MedDT paradoxically reduces vasopressor dosage as lactate escalates, contradicting resuscitation guidelines, while HCT-RL maintains physiologically consistent responses. These findings expose a systemic misalignment between statistical fit and clinical safety, supporting counterfactual audits as a necessary evaluation standard for medical RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。