arXiv:2604.00811stat.MLcs.LG2026-04

提出去混淆得分提升弱重叠场景下的因果效应估计稳定性

Deconfounding Scores and Representation Learning for Causal Effect Estimation with Weak Overlap

  • 引入去混淆得分框架,统一处理倾向性与预后得分
  • 在高维高斯特征下证明预后得分可最优改善重叠性
  • 适合处理特征差异大、重叠弱的因果推断任务

重叠性(positivity)是因果处理效应估计的关键前提。当特征在不同处理组间差异显著时,许多流行估计算法会因方差过高而变得脆弱,尤其在高维情况下,维度诅咒使重叠性难以成立。为此,本文提出一类称为去混淆得分的特征表示,该表示既保持识别性又保留估计目标;经典倾向性得分和预后得分均为其特例。我们将寻找更优重叠性的过程建模为在去混淆得分约束下最小化重叠发散。在广义线性模型与高斯特征的广泛假设下,我们推导出一类去混淆得分的闭式表达,并证明在此类模型中预后得分具有最优重叠性。通过大量实验验证了该性质的实证表现。

原文摘要 · Abstract (English)

Overlap, also known as positivity, is a key condition for causal treatment effect estimation. Many popular estimators suffer from high variance and become brittle when features differ strongly across treatment groups. This is especially challenging in high dimensions: the curse of dimensionality can make overlap implausible. To address this, we propose a class of feature representations called deconfounding scores, which preserve both identification and the target of estimation; the classical propensity and prognostic scores are two special cases. We characterize the problem of finding a representation with better overlap as minimizing an overlap divergence under a deconfounding score constraint. We then derive closed-form expressions for a class of deconfounding scores under a broad family of generalized linear models with Gaussian features and show that prognostic scores are overlap-optimal within this class. We conduct extensive experiments to assess this behavior empirically.

因果推断重叠性表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。