arXiv:2605.22973cs.LGcs.AI2026-05被引 2

提出随机选择作为无监督特征选择的基准,发现许多先进方法反而不如随机选。

Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection

  • 用随机特征选择做基准,评估无监督特征选择方法。
  • 实验表明多数先进方法在性能和效率上均逊于随机选择。
  • 强调新方法必须优于随机基准,否则缺乏实际价值。

每年都有大量新的无监督特征选择方法被提出,但其评估仅限于在选定数据集上计算的有监督与无监督指标,并与已有方法比较。然而,在缺乏明确评估基线的情况下,难以判断这些方法对现有研究的实际贡献及其方法有效性。本文提出使用随机特征选择作为评估无监督特征选择方法的基线。实证结果表明,许多当前最先进的无监督特征选择方法在性能和效率上均被随机特征选择超越。因此,我们强调在开发新型无监督特征选择方法时,必须严格将随机特征选择作为基线,确保其表现始终优于随机选择。

原文摘要 · Abstract (English)

Many novel unsupervised feature selection methods are proposed each year, yet their empirical evaluation is limited to supervised and unsupervised evaluation metrics computed on selected datasets, along with comparisons to existing methods. However, in the absence of an established evaluation baseline, it is difficult to determine the value added to the existing literature by each of these methods, and how effective their underlying approaches are. We propose using random feature selection as a baseline for evaluating the unsupervised feature selection methods. We empirically show that many of the state-of-the-art methods in unsupervised feature selection are outperformed by random feature selection in both performance and efficiency. Accordingly, we emphasize on the strict requirement of considering random feature selection as a baseline in the development process of novel unsupervised feature selection methods to ensure a consistent improvement over random feature selection.

特征选择基准测试无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。