arXiv:2601.15254stat.MLcs.AI2026-01

在环境数据不配对且每环境样本少的情况下,仍能准确估计因果效应。

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

  • 用环境作为工具变量,结合交叉分块的GMM方法解决数据不配对问题。
  • 当环境数增多而每环境样本量固定时,新方法仍能收敛到真实因果效应。
  • 适用于小样本、稀疏效应场景,适合医学或社会科学研究者使用。

我们研究在隐藏混杂因素下的因果效应估计问题,面对的是未配对数据:在不同实验环境(环境)下观测到协变量 $X$ 或结果 $Y$,但无法同时观测两者。在适当正则条件下,该问题可转化为以环境为工具变量(IV)的回归问题。当环境数量多但每个环境观测样本少时,标准两样本IV估计器不再一致。我们提出一种基于工具变量-协变量样本交叉分块的GMM型估计器,并证明其在环境数增长而每环境样本量恒定时具有一致性。进一步通过 $\\(ell_1$-正则化与事后拟合扩展至稀疏因果效应情形。

原文摘要 · Abstract (English)

We study the problem of estimating causal effects under hidden confounding in the following unpaired data setting: we observe some covariates $X$ and an outcome $Y$ under different experimental conditions (environments) but do not observe them jointly; we either observe $X$ or $Y$. Under appropriate regularity conditions, the problem can be cast as an instrumental variable (IV) regression with the environment acting as a (possibly high-dimensional) instrument. When there are many environments but only a few observations per environment, standard two-sample IV estimators fail to be consistent. We propose a GMM-type estimator based on cross-fold sample splitting of the instrument-covariate sample and prove that it is consistent as the number of environments grows but the sample size per environment remains constant. We further extend the method to sparse causal effects via $\ell_1$-regularized estimation and post-selection refitting.

因果推断工具变量小样本稀疏效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。