16种主流因果模型在真实数据上表现普遍不佳,超六成还不如瞎猜。
Do Contemporary Causal Inference Models Capture Real-World Heterogeneity? Findings from a Large-Scale Benchmark
- 用真实世界数据+多种采样策略构建大规模评估基准
- 62%模型的误差比零效应预测还高,80%不如常数效应模型
- 正交性模型仅30%表现最优,质疑其广泛适用性
我们通过大规模基准测试评估了16种现代条件平均处理效应(CATE)估计模型,在12个真实世界数据集上生成43,200个采样变体,采用多样化的观察采样策略。结果发现:(a) 62%的CATE估计误差高于零效应预测器,无效;(b) 即使在存在有效估计的数据集中,仍有80%的模型误差高于常数效应模型;(c) 正交性模型仅在30%情况下表现更优,远低于预期。研究引入新统计量$Q$与$ ilde{Q}$,结合观察采样方法,实现对真实数据中模型排序的无偏估计。所有数据集均来自实地而非模拟,确保反映真实异质性。该基准揭示当前CATE模型在捕捉现实复杂性上的严重不足,亟需更严格的评估框架和方法改进。
原文摘要 · Abstract (English)
We present unexpected findings from a large-scale benchmark study evaluating Conditional Average Treatment Effect (CATE) estimation algorithms, i.e., CATE models. By running 16 modern CATE models on 12 datasets and 43,200 sampled variants generated through diverse observational sampling strategies, we find that: (a) 62\% of CATE estimates have a higher Mean Squared Error (MSE) than a trivial zero-effect predictor, rendering them ineffective; (b) in datasets with at least one useful CATE estimate, 80\% still have higher MSE than a constant-effect model; and (c) Orthogonality-based models outperform other models only 30\% of the time, despite widespread optimism about their performance. These findings highlight significant challenges in current CATE models and underscore the need for broader evaluation and methodological improvements. Our findings stem from a novel application of \textit{observational sampling}, originally developed to evaluate Average Treatment Effect (ATE) estimates from observational methods with experiment data. To adapt observational sampling for CATE evaluation, we introduce a statistical parameter, $Q$, equal to MSE minus a constant and preserves the ranking of models by their MSE. We then derive a family of sample statistics, collectively called $\hat{Q}$, that can be computed from real-world data. When used in observational sampling, $\hat{Q}$ is an unbiased estimator of $Q$ and asymptotically selects the model with the smallest MSE. To ensure the benchmark reflects real-world heterogeneity, we handpick datasets where outcomes come from field rather than simulation. By integrating observational sampling, new statistics, and real-world datasets, the benchmark provides new insights into CATE model performance and reveals gaps in capturing real-world heterogeneity, emphasizing the need for more robust benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。