arXiv:2501.00726math.OCcs.LG2025-01被引 2

通过双重稀疏约束提升无监督特征选择的准确性。

Enhancing Unsupervised Feature Selection via Double Sparsity Constrained Optimization

  • 在经典PCA框架中引入$\ ell_{2,0}$与$\ ell_0$双重稀疏约束
  • 相比最优方法,聚类准确率提升至少3.34%,NMI提升3.02%
  • 适合高维数据中筛选有判别力且去噪的特征子集

无监督特征选择(UFS)广泛应用于机器学习与模式识别。然而,现有方法多仅考虑单一稀疏性,难以从高维特征集中选出有价值且具有判别性的子集。本文提出一种新方法DSCOFS,将双重稀疏约束优化嵌入经典主成分分析(PCA)框架。双重稀疏指同时使用$\ ell_{2,0}$-范数和$\ ell_0$-范数约束变量,通过引入不同类型稀疏性以提升差异特征识别精度。其中,$\ ell_{2,0}$-范数可剔除无关冗余特征,$\ ell_0$-范数能过滤异常噪声特征,二者互补增强判别能力。提出一种有效的近似交替最小化算法求解非凸非光滑模型。理论上严格证明该方法生成序列全局收敛至驻点。在三个合成数据集与八个真实数据集上的数值实验表明,所提方法具有有效性、稳定性和收敛性。尤其在平均聚类准确率(ACC)和归一化互信息(NMI)上,较现有最优方法分别提升至少3.34%和3.02%。更重要的是,两种常见统计检验及新提出的特征相似性度量验证了双重稀疏的优势。所有结果表明,DSCOFS为特征选择提供了新视角。

原文摘要 · Abstract (English)

Unsupervised feature selection (UFS) is widely applied in machine learning and pattern recognition. However, most of the existing methods only consider a single sparsity, which makes it difficult to select valuable and discriminative feature subsets from the original high-dimensional feature set. In this paper, we propose a new UFS method called DSCOFS via embedding double sparsity constrained optimization into the classical principal component analysis (PCA) framework. Double sparsity refers to using $\ell_{2,0}$-norm and $\ell_0$-norm to simultaneously constrain variables, by adding the sparsity of different types, to achieve the purpose of improving the accuracy of identifying differential features. The core is that $\ell_{2,0}$-norm can remove irrelevant and redundant features, while $\ell_0$-norm can filter out irregular noisy features, thereby complementing $\ell_{2,0}$-norm to improve discrimination. An effective proximal alternating minimization method is proposed to solve the resulting nonconvex nonsmooth model. Theoretically, we rigorously prove that the sequence generated by our method globally converges to a stationary point. Numerical experiments on three synthetic datasets and eight real-world datasets demonstrate the effectiveness, stability, and convergence of the proposed method. In particular, the average clustering accuracy (ACC) and normalized mutual information (NMI) are improved by at least 3.34% and 3.02%, respectively, compared with the state-of-the-art methods. More importantly, two common statistical tests and a new feature similarity metric verify the advantages of double sparsity. All results suggest that our proposed DSCOFS provides a new perspective for feature selection.

特征选择稀疏优化无监督学习PCA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。