提出非参数特征选择新方法,可精准控制假发现率且高效识别重要特征。
Nonparametric IPSS: Fast, flexible feature selection with false discovery control
- 基于路径稳定性筛选,适配任意特征重要性评分,无需参数假设。
- 在500样本、5000特征下运行时间<20秒,误报率可控且真阳性检出率更高。
- 适用于高维生物数据,如癌症相关miRNA与基因检测,效果优于现有方法。
特征选择是机器学习与统计中的关键任务。现有方法或依赖线性等参数模型,或缺乏理论上的假发现率控制,或难以识别真正相关的特征。本文提出一种通用的特征选择方法——非参数集成路径稳定性选择(IPSS),其基于任意特征重要性评分,在有限样本下实现假发现率控制。该方法在特征重要性评分非参数时也保持非参数性,并估计更适用于高维数据的q值而非p值。研究聚焦于两种情形:基于梯度提升(IPSSGB)和随机森林(IPSSRF)的重要性评分。在模拟的非线性数据(含RNA测序数据)中,两种方法均准确控制假发现率,并比现有方法检测到更多真阳性特征。两者计算高效,在500样本、5000特征条件下运行时间低于20秒。应用于癌症相关miRNA与基因检测,结果表明其以更少特征获得更好预测性能,优于已有方法。
原文摘要 · Abstract (English)
Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery control, or (iii) identify few true positives. Here, we introduce a general feature selection method with finite-sample false discovery control based on applying integrated path stability selection (IPSS) to arbitrary feature importance scores. The method is nonparametric whenever the importance scores are nonparametric, and it estimates q-values, which are better suited to high-dimensional data than p-values. We focus on two special cases using importance scores from gradient boosting (IPSSGB) and random forests (IPSSRF). Extensive nonlinear simulations with RNA sequencing data show that both methods accurately control the false discovery rate and detect more true positives than existing methods. Both methods are also efficient, running in under 20 seconds when there are 500 samples and 5000 features. We apply IPSSGB and IPSSRF to detect microRNAs and genes related to cancer, finding that they yield better predictions with fewer features than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。