arXiv:2605.10430cs.LGcs.AI2026-05

实证发现:真实评估指标比模拟指标更能反映模型实际效果

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

论文配图:Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation
图 1 · 摘自论文原文
  • 用真实可观察指标与反事实指标双轨评估治疗效应模型
  • 模拟数据上表现好的模型在真实数据中排名不一致
  • 简单元学习器搭配强基模型反而更稳定,适合工业应用

机器学习估计异质性治疗效应受到学术界和工业界的广泛关注。但两者评估方式差异显著:方法研究多依赖半模拟基准与反事实指标,而实际应用则基于可观测的排序或测试结果。尽管方法进展与实际部署之间存在明显差距,但两种评估范式的关系尚未系统探讨。本文对标准半模拟基准族和真实数据集上的治疗效应评估进行了大规模实证研究,涵盖元学习器与多种基础模型,以及专门的因果机器学习模型。我们同时使用应用文献中常见的可观测指标和方法论文中常用的反事实指标进行评估。结果揭示两个互补性差距:其一,反事实指标无法可靠捕捉可观测指标所青睐的估计器,即使在相同半模拟基准上;其二,半模拟基准上的排名无法转移到真实数据集。此外,我们发现简单元学习器搭配强基模型始终表现稳健,优于专用因果模型。总体而言,治疗效应估计的研究进展不应仅通过反事实指标和半模拟基准评估,而应结合可观测指标和真实数据验证。

原文摘要 · Abstract (English)

Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice. However, the two communities often evaluate models under markedly different conditions. Methodological work typically relies on semi-simulated benchmarks and metrics that require counterfactual outcomes, whereas real-world applications rely on observable metrics based on ranking or test outcomes. Despite the well-known gap between methodological progress and practical deployment, the relationship between these evaluation regimes has not been examined systematically. We conduct a large-scale empirical study of treatment effect evaluation across standard semi-simulated benchmark families and real-world datasets. Our benchmark covers meta-learners paired with multiple base learners, as well as specialized causal machine learning models. We evaluate these methods using observable metrics common in application-oriented literature, alongside counterfactual metrics commonly used in methods papers. Our results reveal two complementary gaps. First, counterfactual metrics do not reliably recover the estimators preferred by observable metrics, even on the same semi-simulated benchmarks. Second, rankings obtained on semi-simulated benchmarks do not transfer to real datasets. We further find that simple meta-learners with strong base models are consistently competitive, in contrast to specialized causal models. Overall, our findings suggest that progress in treatment effect estimation research should not be assessed solely through counterfactual metrics and semi-simulated benchmarks, but it would benefit from incorporating observable metrics and real-data validation.

因果推断评估基准真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。