arXiv:2506.13107cs.LGstat.ML2025-06被引 1

诚实估计会降低因果森林的精度,大样本下尤其明显。

Honesty in Causal Forests: When It Helps and When It Hurts

  • 用数据分治法避免过拟合,但限制了对异质效应的捕捉能力
  • 在7000+数据集上,诚实估计需多27%数据才能达到同等效果
  • 应根据实际目标权衡使用,而非默认开启

因果森林用于估计个体间治疗效应的差异,为营销、运营和公共政策等领域的个性化干预提供依据。标准做法是采用‘诚实估计’:将数据分为两部分,一部分用于划分子群体,另一部分用于估计组内治疗效应,以减少过拟合,该方法已是多数软件包的默认设置。然而,这一做法是否总是合适?我们发现,当处理效应异质性显著且数据量足够大时,诚实估计反而会降低个体治疗效应估计的准确性。原因在于偏差-方差权衡:诚实性降低了过拟合风险,却因可用数据减少而增加了欠拟合风险,削弱了对异质性的检测与建模能力。在超过7000个基准数据集上的实验表明,采用默认诚实估计的代价可能高达需额外27%的数据才能达到非诚实模型的性能水平。因此,诚实性本质上是一种正则化手段,是否采用应基于应用目标和实证表现,而非习惯性默认。

原文摘要 · Abstract (English)

Causal forests estimate how treatment effects vary across individuals, guiding personalized interventions in areas like marketing, operations, and public policy. A standard practice is honest estimation: dividing the data into two samples, one to define subgroups and another to estimate treatment effects within them. This is intended to reduce overfitting and is the default in many software packages. But is it the right choice? We show that honest estimation can reduce the accuracy of estimates of individual treatment effects, especially when effect heterogeneity is substantial and datasets are large enough to detect it. The reason is a bias-variance trade-off: honesty lowers the risk of overfitting but increases the risk of underfitting by limiting the data available to detect and model heterogeneity. Across more than 7,000 benchmark datasets, we find that the cost of using honesty by default can be as high as requiring 27% more data to match the performance of models trained without it. Honesty is best understood as a form of regularization. Whether to adopt it should depend on the goals of the application and its empirical performance, not on reflexive default use.

因果推断机器学习统计学习异质性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。