提出修正变量离散化带来的因果效应偏差,提升估计精度。
Coarsening Bias from Variable Discretization in Causal Functionals
- 用组内均值替代离散化后变量,消除主要偏差项
- 在粗粒度分箱下仍实现近似名义置信区间覆盖
- 适用于含隐变量的因果图模型分析
因果识别函数常需对连续变量的条件密度进行积分,如在存在隐藏变量的有向无环图中非参数识别总效应和中介效应时。估计这些密度并计算积分在统计与计算上均具挑战性。常见方法是将连续变量离散化,以有限求和代替积分。尽管简便,离散化会改变总体函数,产生不可忽略的近似偏差,即使识别正确亦然。在光滑性假设下,我们证明该粗化误差为分箱宽度的一阶项,且出现在目标函数层面,区别于统计估计误差。我们提出一种简单去偏的粗化函数,通过在箱内条件均值处评估结果回归,消除主导的粗化误差项,使近似误差降为二阶。推导了该去偏函数的插补估计量与一步估计量。模拟显示,即便在粗分箱下也能显著降低偏差并实现近似名义置信区间覆盖率。研究提供了一个控制变量离散化对参数逼近与统计估计影响的实用框架。
原文摘要 · Abstract (English)
Causal identification functionals often require integration over conditional densities of continuous variables, such as those arising in nonparametric identification theory of total and mediated causal effects in DAGs with hidden variables. Estimating these densities and evaluating the resulting integrals can be statistically and computationally demanding. A common workaround is to discretize the continuous variable and replace integrals with finite sums. Although convenient, discretization alters the population-level functional and can induce non-negligible approximation bias, even when identification is correct. Under smoothness conditions, we show that the resulting coarsening error is first order in the bin width and arises at the level of the target functional, distinct from statistical estimation error. We propose a simple debiased coarsened functional that evaluates the outcome regression at within-bin conditional means, eliminating the leading coarsening error term and yielding a second-order approximation error. We derive plug-in and one-step estimators for this debiased coarsened functional. Simulations demonstrate substantial bias reduction and near-nominal confidence interval coverage, even under coarse binning. Our results provide a simple framework for controlling the impact of variable discretization on both parameter approximation and statistical estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。