剖析全局特征效应估计的误差来源,指导如何选数据和方法。
Analyzing Error Sources in Global Feature Effect Estimation
- 分解误差为模型偏差、估计偏差等四类,分析其影响因素。
- 实证表明用训练数据比验证数据误差更小,因样本量更大。
- 交叉验证能降低过拟合模型的估计方差,适合复杂模型。
局部依赖(PD)和累积局部效应(ALE)等全局特征效应广泛用于解释黑箱模型,但它们仅为真实效应的估计值,其可靠性受多重误差源影响。尽管应用广泛,这些误差来源尚未被系统研究,尤其是应使用训练数据还是预留数据来估计特征效应仍无定论。本文从估计器层面出发,系统分析了PD与ALE的偏差与方差来源,推导出均方误差分解式,分离出模型偏差、估计偏差、模型方差与估计方差,并分析其对模型特性、数据选择与样本量的依赖关系。通过在多种数据生成过程、学习器、估计策略(训练数据、验证数据、交叉验证)及样本量下的大规模模拟实验,验证理论结果。发现虽预留数据理论上更干净,但训练数据带来的偏差在实际中可忽略,且常因更高样本量而整体表现更优;估计方差受交互作用与样本量共同影响,其中ALE对样本量尤为敏感;基于交叉验证的估计可有效降低过拟合模型的模型方差成分。本研究为特征效应估计中的误差来源提供原理性解释,并为模型解释策略选择提供明确指导。
原文摘要 · Abstract (English)
Global feature effects such as partial dependence (PD) and accumulated local effects (ALE) plots are widely used to interpret black-box models. However, they are only estimates of true underlying effects, and their reliability depends on multiple sources of error. Despite the popularity of global feature effects, these error sources are largely unexplored. In particular, the practically relevant question of whether to use training or holdout data to estimate feature effects remains unanswered. We address this gap by providing a systematic, estimator-level analysis that disentangles sources of bias and variance for PD and ALE. To this end, we derive a mean-squared-error decomposition that separates model bias, estimation bias, model variance, and estimation variance, and analyze their dependence on model characteristics, data selection, and sample size. We validate our theoretical findings through an extensive simulation study across multiple data-generating processes, learners, estimation strategies (training data, validation data, and cross-validation), and sample sizes. Our results reveal that, while using holdout data is theoretically the cleanest, potential biases arising from the training data are empirically negligible and dominated by the impact of the usually higher sample size. The estimation variance depends on both the presence of interactions and the sample size, with ALE being particularly sensitive to the latter. Cross-validation-based estimation is a promising approach that reduces the model variance component, particularly for overfitting models. Our analysis provides a principled explanation of the sources of error in feature effect estimates and offers concrete guidance on choosing estimation strategies when interpreting machine learning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。