提出新方法提升因果推断效率,解决传统方法在大规模数据中失效问题。
Efficient Identification of Direct Causal Parents via Invariance and Minimum Error Testing
- 利用误差不等式约束,从潜在变量中筛选直接因果因素。
- 在真实大规模数据集上达到当前最佳性能,显著减少测试次数。
- 适合处理高维、分布变化复杂的因果发现任务。
不变因果预测(ICP)是一种通过利用分布变化和不变性检验来识别目标变量直接因果父母的流行方法(Peters et al., 2016)。然而,由于ICP需要执行指数级数量的检验,并且当分布变化仅影响少数变量时无法识别因果关系,因此在实际大规模问题中应用困难。本文提出MMSE-ICP和fastICP两种方法,利用一个误差不等式解决ICP的可识别性问题。该不等式表明:使用因果父母构建的预测器的最小预测误差,小于所有不使用后代变量的预测器。fastICP是一种面向大规模问题的高效近似方法,它利用该不等式和启发式策略,大幅减少测试次数。在多个模拟实验中,MMSE-ICP和fastICP均优于现有基线方法,并在大规模真实数据基准上取得当前最优结果。
原文摘要 · Abstract (English)
Invariant causal prediction (ICP) is a popular technique for finding causal parents (direct causes) of a target via exploiting distribution shifts and invariance testing (Peters et al., 2016). However, since ICP needs to run an exponential number of tests and fails to identify parents when distribution shifts only affect a few variables, applying ICP to practical large scale problems is challenging. We propose MMSE-ICP and fastICP, two approaches which employ an error inequality to address the identifiability problem of ICP. The inequality states that the minimum prediction error of the predictor using causal parents is the smallest among all predictors which do not use descendants. fastICP is an efficient approximation tailored for large problems as it exploits the inequality and a heuristic to run fewer tests. MMSE-ICP and fastICP not only outperform competitive baselines in many simulations but also achieve state-of-the-art result on a large scale real data benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。