arXiv:2507.17439stat.MEcs.LG2025-07

高维小样本下抗数据污染的因果推断新方法

Penalized Empirical Likelihood for Doubly Robust Causal Inference under Contamination in High Dimensions

  • 结合结果模型与倾向得分平衡,用非凸正则化惩罚经验似然
  • 在部分模型正确时仍保持一致性和渐近正态性,误差更小
  • 适合基因表达等高维数据,对异常值不敏感且区间校准好

针对高维低样本量观测研究中因数据污染和模型误设带来的严重推断挑战,本文提出一种双重稳健的平均处理效应估计器。该方法将有界影响估计方程与协变量平衡倾向得分嵌入非凸正则化的惩罚经验似然框架中,满足极小化性质:在部分模型正确时仍能实现一致性、选择一致性、抗污染能力及渐近正态性。为进行不确定性量化,通过累积量生成函数与影响函数修正构建有限样本置信区间,避免依赖渐近近似。模拟研究及在基因表达数据集(Golub 和 Khan)上的应用表明,该方法在偏差、误差指标和区间校准方面均表现更优,凸显其在高维低样本场景下的稳健性与推断有效性。值得注意的是,即使无污染情况下,该估计器及其置信区间也比现有方法更高效。

原文摘要 · Abstract (English)

We propose a doubly robust estimator for the average treatment effect in high dimensional low sample size observational studies, where contamination and model misspecification pose serious inferential challenges. The estimator combines bounded influence estimating equations for outcome modeling with covariate balancing propensity scores for treatment assignment, embedded within a penalized empirical likelihood framework using nonconvex regularization. It satisfies the oracle property by jointly achieving consistency under partial model correct ness, selection consistency, robustness to contamination, and asymptotic normality. For uncertainty quantification, we derive a finite sample confidence interval using cumulant generating functions and influence function corrections, avoiding reliance on asymptotic approximations. Simulation studies and applications to gene expression datasets (Golub and Khan) demonstrate superior performance in bias, error metrics, and interval calibration, highlighting the method robustness and inferential validity in HDLSS regimes. One notable aspect is that even in the absence of contamination, the proposed estimator and its confidence interval remain efficient compared to those of competing models.

因果推断高维数据稳健估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。