提出新方法A2D2E,更稳定地估计黑箱模型主效应。
Accumulated Aggregated D-Optimal Designs for Estimating Main Effects in Black-Box Models
- 将主效应估计视为设计问题,用D-optimal超立方体设计替代传统采样点。
- 在高特征相关下表现显著优于ALE方法,误差降低超过30%。
- 适用于无梯度要求的黑箱模型,适合可解释性研究者使用。
在可解释机器学习中,估算输入变量对黑箱模型输出的影响是核心任务。现有方法存在两大缺陷:在分布外(OOD)评估时敏感,当查询点远离数据流形时效果下降;在特征相关下不稳定,导致实际估计不可靠。本文将主效应估计统一建模为设计问题,揭示现有方法仅在评估位置选择上不同。基于此,提出A2D2E——基于累积聚合D-最优设计的估计器,用D-optimal超立方体设计替代传统采样,最小化主效应估计方差。A2D2E具有模型无关性,无需预测器可微,且有闭式解,计算复杂度与现有方法相当。理论证明其一致收敛于与ALE相同的总体目标,并扩展至仅有代理模型可用的现实场景。多模型、多相关结构下的大量模拟表明,A2D2E显著优于基于ALE的方法,尤其在高特征相关下优势明显。
原文摘要 · Abstract (English)
Estimating how individual input variables affect the output of a black-box model is a central task in explainable machine learning. However, existing methods suffer from two key limitations: sensitivity to out-of-distribution (OOD) evaluations, which arises when query points are placed far from the data manifold, and instability under feature correlation, which can lead to unreliable effect estimates in practice. We introduce a unified view of main effect estimation as a design problem, which reveals that all existing methods differ only in their choice of evaluation locations. Building on this formulation, we propose A2D2E, an Estimator based on Accumulated Aggregated D-Optimal Designs, which replaces evaluations with a D-optimal hypercube design to minimize the variance of main effect estimation. A2D2E is model-agnostic, requires no differentiability of the predictor, and admits a closed-form estimator with complexity comparable to existing approaches. We establish that A2D2E is consistent to the same population target as ALE, and extend this result to the realistic setting where only a surrogate model is available. Through extensive simulations across multiple predictive models and dependence settings, we demonstrate that A2D2E outperforms ALE-based methods, with the largest gains under high feature correlation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。