用双重稳健方法学习低成本序列检测策略,提升医疗决策效率。
Cost-optimal Sequential Testing via Doubly Robust Q-learning
- 引入路径特异性加权,处理测试结果依赖的缺失数据问题。
- 在任一模型正确时仍可无偏估计,实现成本优化策略学习。
- 适用于需权衡检测成本与准确性的临床研究场景。
临床决策常涉及昂贵、侵入性或耗时的检测,促使制定个性化、序贯的检测时机与终止策略。本文研究从回顾性数据中学习成本最优的序贯决策策略,其中检测可用性依赖于前期结果,导致信息性缺失。在序贯缺失随机机制下,提出一种双重稳健Q-learning框架。该方法引入路径特异性逆概率权重,以处理异质检测轨迹,并满足观测历史条件下的归一化性质。结合辅助对比模型,构建正交伪输出,使得当获取模型或对比模型之一正确时,政策学习仍无偏。建立了阶段式对比估计器的极值不等式、收敛速率、策略后悔界及误分类率。模拟实验显示,相比加权和完整案例基线,该方法在成本调整性能上表现更优;在前列腺癌队列研究中的应用表明,该方法可在不降低预测准确性前提下减少检测成本。
原文摘要 · Abstract (English)
Clinical decision-making often involves selecting tests that are costly, invasive, or time-consuming, motivating individualized, sequential strategies for what to measure and when to stop ascertaining. We study the problem of learning cost-optimal sequential decision policies from retrospective data, where test availability depends on prior results, inducing informative missingness. Under a sequential missing-at-random mechanism, we develop a doubly robust Q-learning framework for estimating optimal policies. The method introduces path-specific inverse probability weights that account for heterogeneous test trajectories and satisfy a normalization property conditional on the observed history. By combining these weights with auxiliary contrast models, we construct orthogonal pseudo-outcomes that enable unbiased policy learning when either the acquisition model or the contrast model is correctly specified. We establish oracle inequalities for the stage-wise contrast estimators, along with convergence rates, regret bounds, and misclassification rates for the learned policy. Simulations demonstrate improved cost-adjusted performance over weighted and complete-case baselines, and an application to a prostate cancer cohort study illustrates how the method reduces testing cost without compromising predictive accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。