arXiv:2502.00501stat.MEcs.AI2025-02被引 1

提出三阶段框架优化因果推断特征选择,降低估计偏差与方差。

Optimizing Feature Selection in Causal Inference: A Three-Stage Computational Framework for Unbiased Estimation

  • 分三阶段筛选关键变量,避免引入混淆或无关特征
  • 在合成数据上显著降低因果估计偏差与方差
  • 适用于大规模真实数据,如美国阿片类药物危机研究

特征选择在因果推断中至关重要,直接影响因果效应估计的无偏性。恰当的特征可缩短匹配算法耗时,并有效降低估计偏差与方差。现有方法常因仅平衡处理相关变量导致偏差,或过度平衡虚假变量增加方差。本文提出一种改进的三阶段计算框架,在多种设置下使用先进合成数据评估,相比当前最优方法显著提升变量筛选效果,实现更低的偏差与方差,且计算效率可行,具备大规模数据扩展性。为验证实际应用价值,进一步应用于美国阿片类药物滥用与自杀行为之间的因果关系分析。

原文摘要 · Abstract (English)

Feature selection is an important but challenging task in causal inference for obtaining unbiased estimates of causal quantities. Properly selected features in causal inference not only significantly reduce the time required to implement a matching algorithm but, more importantly, can also reduce the bias and variance when estimating causal quantities. When feature selection techniques are applied in causal inference, the crucial criterion is to select variables that, when used for matching, can achieve an unbiased and robust estimation of causal quantities. Recent research suggests that balancing only on treatment-associated variables introduces bias while balancing on spurious variables increases variance. To address this issue, we propose an enhanced three-stage framework that shows a significant improvement in selecting the desired subset of variables compared to the existing state-of-the-art feature selection framework for causal inference, resulting in lower bias and variance in estimating the causal quantity. We evaluated our proposed framework using a state-of-the-art synthetic data across various settings and observed superior performance within a feasible computation time, ensuring scalability for large-scale datasets. Finally, to demonstrate the applicability of our proposed methodology using large-scale real-world data, we evaluated an important US healthcare policy related to the opioid epidemic crisis: whether opioid use disorder has a causal relationship with suicidal behavior.

因果推断特征选择偏差控制医疗政策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。