arXiv:2605.13430stat.MEcs.AI2026-05

揭示选择偏差下因果效应可识别的充要条件,为真实数据中的因果推断提供理论保障。

Towards a holistic understanding of Selection Bias for Causal Effect Identification

论文配图:Towards a holistic understanding of Selection Bias for Causal Effect Identification
图 1 · 摘自论文原文
  • 基于弱假设刻画倾向得分与选择概率,推导出平均处理效应可识别的充要条件。
  • 在存在选择偏差时,所提条件比已有方法更宽松且覆盖更广。
  • 适用于生物银行等存在健康志愿者偏差的真实观测数据,适合因果推断研究者阅读。

选择偏差广泛存在于观察性研究中。例如,大规模生物银行数据可能因参与者更健康、社会经济地位更高而产生‘健康志愿者偏差’。从这类子人群恢复总体的平均处理效应(ATE)是因果推断中的关键问题,因为仅基于选中群体估计的 ATE 可能严重偏离总体真实值。本文研究在选择偏差下 ATE 的可识别性,提出必要且充分的识别条件,通过在概率类上施加弱假设来刻画倾向得分和选择概率。相比以往工作,我们的结果扩展了现有的图模型可识别性准则,在相同条件下具有更严格的适用性,提供了对选择偏差下因果效应识别的更全面理解。

原文摘要 · Abstract (English)

Selection bias is pervasive in observational studies. For example, large scale biobanks data can exhibit ``healthy volunteer bias'' when respondents are healthier and of higher socio-economic status than the population they are meant to represent. Recovering causal effects from such sub-population is an important problem in causal inference, as estimating average treatment effects (ATE) from selected populations can result in a severely biased estimate of the ATE from the whole population. In this paper, we investigate the identifiability of the ATE under selection bias. We provide necessary and sufficient conditions for ATE identifiability, leveraging weak assumptions on probability classes to characterize propensity score and selection probability. Compared to previous works, our results extend existing graphical identifiability criteria and offer a more comprehensive understanding of causal effect identification with strictly weaker conditions in the presence of selection bias.

因果推断选择偏差可识别性平均处理效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。