arXiv:2511.08303stat.MLcs.LG2025-11

利用未标注协变量提升因果效应估计效率

Semi-Supervised Treatment Effect Estimation with Unlabeled Covariates for Prediction-Powered Causal Inference

  • 基于半监督学习框架,融合未标注协变量信息
  • 在两种数据设置下均降低估计方差,提升效率
  • 适合有大量未标注数据的医学、社会科学研究

本研究探讨半监督环境下因果效应估计问题,亦可理解为预测驱动的因果推断。在该设定中,除标准的协变量、处理指示变量和结果变量外,还可利用未标注的辅助协变量。针对此问题,我们推导了效率界限,并提出了渐近方差达到效率界限的高效估计器。分析中引入两种数据生成过程:单样本设置(即部分数据可观测处理指标与结果,也称删失设置)和双样本设置(即独立的标记与未标记数据集,又称病例对照或分层设置)。在两种情形下,通过引入辅助协变量,均可降低效率界限,并获得比不使用辅助协变量时具有更小渐近方差的估计器。我们将该框架定义为预测驱动的因果推断。

原文摘要 · Abstract (English)

This study investigates treatment effect estimation in the semi-supervised setting, also can be interpreted as prediction-powered inference. In our setting, we can use not only the standard triple of covariates, treatment indicator, and outcome, but also unlabeled auxiliary covariates. For this problem, we develop efficiency bounds and efficient estimators whose asymptotic variance aligns with the efficiency bound. In the analysis, we introduce two different data-generating processes: the one-sample setting and the two-sample setting. The one-sample setting considers the case where we can observe treatment indicators and outcomes for a part of the dataset, which is also called the censoring setting. In contrast, the two-sample setting considers two independent datasets with labeled and unlabeled data, which is also called the case-control setting or the stratified setting. In both settings, we find that by incorporating auxiliary covariates, we can lower the efficiency bound and obtain an estimator with an asymptotic variance smaller than that without such auxiliary covariates. We frame our framework as prediction-powered causal inference.

因果推断半监督学习预测驱动效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。