arXiv:2605.24808cs.LGcs.AI2026-05

提出新方法提升因果效应估计精度,解决高维数据偏差问题。

Disentangled Double Machine Learning for Accurate Causal Effect Estimation

论文配图:Disentangled Double Machine Learning for Accurate Causal Effect Estimation
图 1 · 摘自论文原文
  • 分离混杂因子、处理特异与结果特异因素,提升干扰项估计可靠性。
  • 在合成与真实数据上,平均绝对误差和均方根误差均显著优于13种基线。
  • 适合处理高维观测数据中的因果推断,尤其关注估计稳定性与准确性。

混淆偏差是基于观测数据进行因果效应估计的关键挑战。双机器学习(DML)通过估计处理和结果的干扰函数,构建残差并从残差中估计因果效应来应对该问题。然而,在高维或小样本情形下,DML常产生有偏且不稳定的估计结果。原因之一是其使用全部协变量估计干扰函数,未区分潜在因子,导致干扰估计不可靠;另一原因是干扰估计不准确会引入处理残差与剩余结果误差之间的残差依赖,影响因果效应估计的准确性。为此,本文提出解耦双机器学习(DDML),融合两项关键策略:首先,采用因果角色解耦策略将协变量分解为混杂因子、处理特异因子和结果特异因子,以实现可靠的干扰函数估计;其次,采用残差依赖正交化策略缓解由干扰估计误差引起的残差依赖,提升因果效应估计精度。在合成、半合成及真实世界数据集上的实验表明,DDML在平均绝对误差(MAE)和均方根误差(RMSE)上显著优于13种先进基线算法。

原文摘要 · Abstract (English)

Confounding bias is a key challenge in causal effect estimation from observational data. Double Machine Learning (DML) addresses this issue by estimating treatment and outcome nuisance functions, constructing treatment and outcome residuals, and estimating causal effects from the residuals. However, DML often produces biased and unstable estimates in highdimensional or finite-sample scenarios. One reason is that DML estimates nuisance functions using all covariates without disentangling distinct latent factors, resulting in unreliable nuisance function estimation. Another is that imprecise nuisance estimation further introduces residual dependence between the treatment residual and the remaining outcome error, undermining the accuracy of causal effect estimates. To address these issues, in this paper, we propose Disentangled Double Machine Learning (DDML), a novel algorithm that integrates two key strategies. First, a causal role disentanglement strategy decomposes covariates into confounders, treatment-specific factors, and outcomespecific factors for enabling reliable nuisance function estimation. And second, a residual dependence orthogonalization strategy mitigates residual dependence caused by nuisance estimation errors for enhancing the precision of causal effect estimates. Experimental results on synthetic, semi-synthetic, and real-world datasets demonstrate that DDML significantly outperforms 13 state-of-the-art baseline algorithms in both MAE and RMSE.

因果推断机器学习估计精度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。