arXiv:2409.09003stat.MLcs.LG2024-09被引 4

不依赖模型的变量选择新方法,无需生成假数据即可高效筛选关键特征。

Model-independent variable selection via the rule-based variable priority

  • 基于规则计算样本统计量平均值,实现无需预测误差评估的变量优先级排序。
  • 理论证明对噪声变量具有一致筛选能力,实测在多种数据类型中表现稳定。
  • 适用于回归、分类与生存分析,适合追求可解释性的研究者使用。

尽管高预测精度是机器学习的基本目标,但识别少数具有高解释力的特征同样重要。传统方法如置换重要性通过扰动变量并测量预测误差变化来评估其影响,但需生成人工数据,且多为模型特定。本文提出一种新型模型无关方法——变量优先级(VarPro),利用规则直接计算样本简单统计量的平均值,无需生成人工数据或评估预测误差。该方法易于操作,适用于回归、分类和生存分析等多种场景。我们研究了VarPro的渐近性质,证明其对噪声变量具有稳定的筛选一致性。在合成数据与真实数据上的实证研究表明,该方法性能均衡,优于多数当前主流变量选择方法。

原文摘要 · Abstract (English)

While achieving high prediction accuracy is a fundamental goal in machine learning, an equally important task is finding a small number of features with high explanatory power. One popular selection technique is permutation importance, which assesses a variable's impact by measuring the change in prediction error after permuting the variable. However, this can be problematic due to the need to create artificial data, a problem shared by other methods as well. Another problem is that variable selection methods can be limited by being model-specific. We introduce a new model-independent approach, Variable Priority (VarPro), which works by utilizing rules without the need to generate artificial data or evaluate prediction error. The method is relatively easy to use, requiring only the calculation of sample averages of simple statistics, and can be applied to many data settings, including regression, classification, and survival. We investigate the asymptotic properties of VarPro and show, among other things, that VarPro has a consistent filtering property for noise variables. Empirical studies using synthetic and real-world data show the method achieves a balanced performance and compares favorably to many state-of-the-art procedures currently used for variable selection.

变量选择模型无关可解释性统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。