arXiv:2506.10929stat.MEcs.LG2025-06被引 1

针对双不平衡数据,提出基于最小深度的稳定特征选择方法。

On feature selection in double-imbalanced data settings: a Random Forest approach

  • 利用树结构拓扑信息评估变量重要性,改进传统随机森林特征选择。
  • 在模拟与真实数据上,选出的特征子集更简洁且分类更准确。
  • 适合处理高维小样本且类别不均衡的数据场景,结果可解释性强。

在高维分类任务中,特征选择至关重要,尤其在双不平衡场景下——即响应变量存在类别不平衡,同时数据呈现维度不对称(n ≫ p)。此类情况下,传统应用于随机森林(RF)的特征选择方法常导致不稳定或误导性的重要性排序。本文提出一种基于最小深度的新阈值筛选方案,利用树结构拓扑评估变量相关性。在模拟与真实数据集上的大量实验表明,该方法生成的特征子集比传统最小深度法更简洁、更准确。该方法为随机森林在双不平衡条件下的变量选择提供了实用且可解释的解决方案。

原文摘要 · Abstract (English)

Feature selection is a critical step in high-dimensional classification tasks, particularly under challenging conditions of double imbalance, namely settings characterized by both class imbalance in the response variable and dimensional asymmetry in the data $(n \gg p)$. In such scenarios, traditional feature selection methods applied to Random Forests (RF) often yield unstable or misleading importance rankings. This paper proposes a novel thresholding scheme for feature selection based on minimal depth, which exploits the tree topology to assess variable relevance. Extensive experiments on simulated and real-world datasets demonstrate that the proposed approach produces more parsimonious and accurate subsets of variables compared to conventional minimal depth-based selection. The method provides a practical and interpretable solution for variable selection in RF under double imbalance conditions.

特征选择随机森林双不平衡高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。