Offline RL模型在中风治疗中的优越性可能源于数据混淆,而非真实疗效。
Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry

- 通过多算法、多奖励设计的系统评估,发现模型提升效果由数据偏差导致
- 去除混淆因素后,模型优势消失且不显著(p > 0.1)
- 提出六步评估清单,适用于医疗决策类强化学习研究
近期离线强化学习研究声称其策略优于医生决策。本研究对五大离线RL算法家族与14种奖励设计,在来自全国注册库的44,894名2018年后急性缺血性中风患者(总计129,033例)中进行系统性部分交叉评估。标准拟合Q评估(FQE)显示政策改进估计值为+0.0069;加入早期神经功能恶化惩罚后增至+0.0101。我们识别出奖励嵌入混淆:终端奖励同时编码了基线严重程度、预后及治疗效果。2×2因子分析表明,终端奖励混淆解释了218.6%的观察信号变化,移除后反而超过零假设。经基于双重机器学习(DML)思想的广义可加模型(GBM)奖励残差化处理,FQE估计降至+0.0033(p = 0.132),完全去混淆后为+0.0025(p = 0.291)。FQE诊断、T-learner分析与直接重复分析均指向无临床意义的整体改善。1年改良秩次量表(mRS)分析复现该衰减。本文提供一个基于实证的六步评估检查清单。NIHSS分层异质性为前瞻性试验设计提供假说生成依据;医院间分歧在完全去混淆后不再持续。
原文摘要 · Abstract (English)
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). Standard Fitted Q-Evaluation (FQE) yields an apparent policy-improvement estimate of +0.0069; adding an Early Neurological Deterioration penalty increases it to +0.0101. We identify reward-embedded confounding, in which a proxy terminal reward encodes baseline severity and prognosis as well as treatment efficacy. A 2 x 2 factorial analysis finds that terminal reward confounding accounts for 218.6% of the observed signal change, so its removal overshoots the null. After DML-inspired GBM reward residualization, the FQE estimate attenuates to +0.0033 (p = 0.132), and full deconfounding yields +0.0025 (p = 0.291). FQE-based diagnostics, T-learner analyses, and direct recurrence analyses converge away from a clinically meaningful aggregate improvement. A 1-year mRS factorial analysis replicates the attenuation. We provide an empirically motivated six-step evaluation checklist. NIHSS-stratified heterogeneity is hypothesis-generating for prospective trial design; hospital-level disagreement does not persist after full reward deconfounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。