联合估计多个治疗策略,显著降低预测误差。
Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings

- 设计共享策略编码器,实现多策略联合建模
- 在模拟数据上根均方误差大幅下降
- 适合评估相似治疗方案的医疗决策场景
在纵向研究中,比较多个动态治疗策略对医疗和政策决策至关重要,但传统方法对每个策略单独估计,无法共享反事实信息。我们发现这种独立估计会引入结构上未控制的二阶偏差,即使使用纵向目标最大似然估计(LTMLE)修正后,有限样本方差仍被放大。为此,我们提出一种策略感知的迭代条件期望(ICE)Q函数重参数化方法,通过共享表示实现联合估计。我们构建了策略编码的Q网络(PEQ-Net),其核心为共享策略编码器,采用核均值嵌入训练,使学习到的表征空间反映群体层面的策略差异。经LTMLE校正后,该设计对二阶余项施加结构约束,从而稳定有限样本方差。在半合成数据集上的实验表明,PEQ-Net持续优于现有基于ICE的方法,在评估相近策略时,根均方误差显著降低。
原文摘要 · Abstract (English)
Comparative evaluation of multiple dynamic treatment policies is essential for healthcare and policy decisions, yet conventional longitudinal causal inference methods estimate each in isolation, preventing information sharing across counterfactuals. We demonstrate that this separate estimation paradigm induces a structurally uncontrolled second-order bias, inflating finite-sample variance even after standard debiasing with longitudinal targeted maximum likelihood estimation(LTMLE). To address this, we propose a policy-aware reparameterization of Iterative Conditional Expectation (ICE) Q-functions that enables joint estimation through shared representations. We implement this approach in the Policy-Encoded Q Network (PEQ-Net), an architecture centered on a shared policy encoder. The encoder is trained using kernel mean embeddings, ensuring that the learned representation space reflects population-level policy dissimilarities. After applying an LTMLE correction step, we prove this design imposes a structural constraint on the second-order remainder, thereby stabilizing finite-sample variance. Experiments on semi-synthetic datasets demonstrate that PEQ-Net consistently outperforms existing ICE-based methods, achieving substantial reductions in root-mean-square error, particularly when evaluating closely related policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。