比较无模型依赖特征选择方法的效率,发现GCM优于LOCO。
Comparing Model-agnostic Feature Selection Methods through Relative Efficiency
- 用相对变异性衡量不同统计量的效率,构建通用比较框架。
- 理论与实证均显示,在适定条件下GCM方法表现更优。
- 适用于神经网络、梯度提升树等常见机器学习模型。
在无模型依赖的特征选择与重要性估计中,包装法因具备通用性而被广泛使用。本文基于相对效率构建通用比较框架,引入相对变异性 $σ/μ$ 以校正不同统计量均值差异的影响。重点对比当前先进方法:广义协方差度量(GCM)与留一协变量法(LOCO)。理论分析涵盖线性模型、非线性可加模型及模拟单层神经网络的单指标模型。结合模拟实验与真实数据案例,验证三种模型及模型误设情形下的表现。结果表明,在由相关性量度定义的合适正则条件下,GCM相关方法在渐近相对效率上普遍优于LOCO。实验包含神经网络与梯度提升树等主流机器学习方法。
原文摘要 · Abstract (English)
Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest. Wrapper methods are commonly used because they are typically model-agnostic. In this paper, we develop a general comparison framework for model-agnostic feature selection methods based on relative efficiency, using \emph{relative variability} $σ/μ$ to account for different statistics having different means. In particular we focus on state-of-the-art feature selection methods, the Generalized Covariance Measure (GCM) and Leave-One-Covariate-Out (LOCO) estimation. In particular, we present a theoretical comparison under three model settings: linear models, non-linear additive models, and single index models that mimic a single-layer neural network. We complement this with simulations and real data examples for the above models and mis-specified models. Our theoretical results, along with empirical findings, demonstrate that GCM-related methods generally out-perform LOCO under suitable regularity conditions defined by a suitably defined correlation quantity which quantifies the asymptotic relative efficiency of these approaches. Our simulations and real data analysis include widely used machine learning methods such as neural networks and gradient boosting trees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。