随机选0.02%-1%特征就能媲美全集,质疑特征选择的科学意义
On the (In)Significance of Feature Selection in High-Dimensional Datasets
- 从高维数据中随机抽取0.02%-1%特征,性能与全集相当
- 在30个数据集中,28个中随机特征表现不输甚至优于特征选择结果
- 提醒需严格验证,避免误将随机结果当有意义信号
特征选择(FS)通常被认为能提升预测性能并识别有意义特征。然而,在30个多样数据集(包括微阵列、批量和单细胞RNA-Seq、质谱、影像等)中,仅0.02%-1%的随机特征子集在28个数据集中表现与全特征集或经计算筛选的特征集相当甚至更优。这表明任意特征集合的表现差异极小,因此若随机特征也能达到同样效果,被选出的特征何以称其为‘重要’?研究挑战了特征选择可靠捕捉有效信号的假设,强调在计算基因组学中,必须经过严格验证才能将所选特征解释为可操作的生物学信号。
原文摘要 · Abstract (English)
Feature selection (FS) is assumed to improve predictive performance and identify meaningful features in high-dimensional datasets. Surprisingly, small random subsets of features (0.02-1%) match or outperform the predictive performance of both full feature sets and FS across 28 out of 30 diverse datasets (microarray, bulk and single-cell RNA-Seq, mass spectrometry, imaging, etc.). In short, any arbitrary set of features is as good as any other (with surprisingly low variance in results) - so how can a particular set of selected features be "important" if they perform no better than an arbitrary set? These results challenge the assumption that computationally selected features reliably capture meaningful signals, emphasizing the importance of rigorous validation before interpreting selected features as actionable, particularly in computational genomics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。