arXiv:2601.08257cs.LGcs.AI2026-01

提出多标签评估框架,更公平地比较无监督特征选择方法

On Evaluation of Unsupervised Feature Selection for Pattern Classification

  • 用多标签数据替代单标签数据评估特征选择效果
  • 21个数据集实验显示排名与传统方法差异显著
  • 适合关注特征选择公平性与可复现性的研究者

无监督特征选择旨在不依赖标签的情况下选出能捕捉数据内在结构的紧凑特征子集。现有研究大多使用从多标签数据中随机选取一个标签构建的单标签数据集进行评估,但所选标签可能因实验设置不同而随意变化,导致方法优劣判断不稳定。本文重新审视这一评估范式,引入多标签分类框架。在21个多标签数据集上对多个代表性方法进行实验,结果表明性能排名与单标签设定下存在明显差异,说明多标签评估能提供更公平、可靠的无监督特征选择方法比较基础。

原文摘要 · Abstract (English)

Unsupervised feature selection aims to identify a compact subset of features that captures the intrinsic structure of data without supervised label. Most existing studies evaluate the performance of methods using the single-label dataset that can be instantiated by selecting a label from multi-label data while maintaining the original features. Because the chosen label can vary arbitrarily depending on the experimental setting, the superiority among compared methods can be changed with regard to which label happens to be selected. Thus, evaluating unsupervised feature selection methods based solely on single-label accuracy is unreasonable for assessing their true discriminative ability. This study revisits this evaluation paradigm by adopting a multi-label classification framework. Experiments on 21 multi-label datasets using several representative methods demonstrate that performance rankings differ markedly from those reported under single-label settings, suggesting the possibility of multi-label evaluation settings for fair and reliable comparison of unsupervised feature selection methods.

特征选择多标签评估方法无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。