arXiv:2604.11165stat.MLcs.AI2026-04

用双重稳健方法学习低成本序列检测策略,提升医疗决策效率。

Cost-optimal Sequential Testing via Doubly Robust Q-learning

  • 引入路径特异性加权,处理测试结果依赖的缺失数据问题。
  • 在任一模型正确时仍可无偏估计,实现成本优化策略学习。
  • 适用于需权衡检测成本与准确性的临床研究场景。

临床决策常涉及昂贵、侵入性或耗时的检测,促使制定个性化、序贯的检测时机与终止策略。本文研究从回顾性数据中学习成本最优的序贯决策策略,其中检测可用性依赖于前期结果,导致信息性缺失。在序贯缺失随机机制下,提出一种双重稳健Q-learning框架。该方法引入路径特异性逆概率权重,以处理异质检测轨迹,并满足观测历史条件下的归一化性质。结合辅助对比模型,构建正交伪输出,使得当获取模型或对比模型之一正确时,政策学习仍无偏。建立了阶段式对比估计器的极值不等式、收敛速率、策略后悔界及误分类率。模拟实验显示,相比加权和完整案例基线,该方法在成本调整性能上表现更优;在前列腺癌队列研究中的应用表明,该方法可在不降低预测准确性前提下减少检测成本。

原文摘要 · Abstract (English)

Clinical decision-making often involves selecting tests that are costly, invasive, or time-consuming, motivating individualized, sequential strategies for what to measure and when to stop ascertaining. We study the problem of learning cost-optimal sequential decision policies from retrospective data, where test availability depends on prior results, inducing informative missingness. Under a sequential missing-at-random mechanism, we develop a doubly robust Q-learning framework for estimating optimal policies. The method introduces path-specific inverse probability weights that account for heterogeneous test trajectories and satisfy a normalization property conditional on the observed history. By combining these weights with auxiliary contrast models, we construct orthogonal pseudo-outcomes that enable unbiased policy learning when either the acquisition model or the contrast model is correctly specified. We establish oracle inequalities for the stage-wise contrast estimators, along with convergence rates, regret bounds, and misclassification rates for the learned policy. Simulations demonstrate improved cost-adjusted performance over weighted and complete-case baselines, and an application to a prostate cancer cohort study illustrates how the method reduces testing cost without compromising predictive accuracy.

序贯决策医疗优化双重稳健缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。