arXiv:2606.21185stat.MLcs.LG2026-06

揭示因果推断中隐藏的双重不稳定性,助你选对估算方法

Two Layers of Instability in Causal Estimation

  • 将标准估计量视为结构因果模型的多峰分布点摘要
  • 逆倾向得分与回归估计量存在数据分布跳跃性不连续
  • 后验均值/中位数更稳定,适合对稳定性要求高的场景

从观察数据中进行因果推断具有内在困难,即便在可识别性成立的前提下。Robins和Ritov(1997)及Robins等(2003)表明,因果效应可能是数据分布的不连续函数:任意接近的两个数据分布可能对应不同的因果效应。这一现象独立于估计器选择;但并非所有估计器稳定性相同。本文揭示了依赖于估计器选择的第二层不稳定性。我们证明,许多标准点估计可被看作结构因果模型空间上多峰分布的点摘要,导致估计量随数据分布发生不连续跳跃。该发现构建了一种估计器分类体系,具有决策论意义:稳定性取决于估计器所隐含优化的损失函数是否与因果效应本身一致。具体而言,逆倾向得分估计量和回归估计量属于不连续摘要,而显式后验均值和中位数则为连续估计。

原文摘要 · Abstract (English)

There is a precise sense in which drawing causal inferences from observational data is hard, even when identifiability is assumed. In particular, Robins and Ritov (1997) and Robins et al. (2003) showed that causal effects can be discontinuous as a function of the data distribution: two arbitrarily close data distributions might correspond to different causal effects. This is a fact independent of the choice of estimator; however, not all estimators are equally unstable. Our contribution is to surface a second layer of instability that depends on the choice of estimator. We show that many standard point estimates can be read as point summaries of multimodal distributions over the space of structural causal models. As such, estimators can jump discontinuously in the data distribution. This defines a taxonomy of estimators that admits a decision-theoretic reading: stability depends on whether the implicit loss function an estimator optimizes is aligned with the causal effect itself. Specifically, inverse propensity weighted estimators and regression estimators are examples of discontinuous summaries, while explicit posterior means and medians are shown to be continuous.

因果推断估计稳定性结构模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。