arXiv:2506.12304cs.LGstat.ML2025-06

利用小规模随机试验数据纠正观测数据中的隐藏混杂偏误

Conditional Average Treatment Effect Estimation Under Hidden Confounders

  • 用伪混杂因子生成器融合随机试验与观测数据
  • 在真实和合成数据上显著降低因果估计偏差
  • 适用于隐私敏感场景,无需试验组协变量信息

估计条件潜在结果和条件平均处理效应(CATE)的主要挑战之一是隐藏混杂因素的存在。由于仅凭观测数据无法检验隐藏混杂因素,现有文献普遍假设条件无混杂性。然而在此假设下,未观测混杂因素仍可能导致CATE估计严重偏误。本文考虑一种情形:除大规模观测数据外,还可用一个小规模随机对照试验(RCT)数据集。值得注意的是,我们不对RCT数据集的协变量信息作任何假设,仅要求可观测结果。我们提出一种基于伪混杂因子生成器与对齐潜在结果的CATE估计方法,使观测数据学习到的潜在结果与来自RCT的真实结果对齐。该方法适用于多种实际场景,尤其在隐私敏感领域(如医疗应用)具有优势。通过大量数值实验,证明了该方法在合成与真实数据集上的有效性。

原文摘要 · Abstract (English)

One of the major challenges in estimating conditional potential outcomes and conditional average treatment effects (CATE) is the presence of hidden confounders. Since testing for hidden confounders cannot be accomplished only with observational data, conditional unconfoundedness is commonly assumed in the literature of CATE estimation. Nevertheless, under this assumption, CATE estimation can be significantly biased due to the effects of unobserved confounders. In this work, we consider the case where in addition to a potentially large observational dataset, a small dataset from a randomized controlled trial (RCT) is available. Notably, we make no assumptions on the existence of any covariate information for the RCT dataset, we only require the outcomes to be observed. We propose a CATE estimation method based on a pseudo-confounder generator and a CATE model that aligns the learned potential outcomes from the observational data with those observed from the RCT. Our method is applicable to many practical scenarios of interest, particularly those where privacy is a concern (e.g., medical applications). Extensive numerical experiments are provided demonstrating the effectiveness of our approach for both synthetic and real-world datasets.

因果推断隐藏混杂随机试验潜在结果

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。