arXiv:2512.18720stat.MLcs.LG2025-12

用深度自编码器提升无监督特征选择的非线性建模与抗噪能力

Unsupervised Feature Selection via Robust Autoencoder and Adaptive Graph Learning

  • 通过深度自编码器学习非线性特征表示,替代传统线性投影
  • 在含异常值数据上性能优于现有最优方法,准确率显著提升
  • 适合高维数据中存在噪声或异常点的特征选择场景

有效的特征选择对高维数据分析和机器学习至关重要。无监督特征选择(UFS)旨在同时完成数据聚类并识别最具区分性的特征。现有UFS方法通常将特征线性映射到伪标签空间进行聚类,但存在两个关键缺陷:(1) 过于简化的线性映射难以捕捉复杂特征关系;(2) 假设簇分布均匀,忽略真实数据中普遍存在的异常值。为此,我们提出基于鲁棒自编码器的无监督特征选择模型(RAEUFS),利用深度自编码器学习非线性特征表示,同时天然增强对异常值的鲁棒性。我们还设计了高效的优化算法。大量实验表明,该方法在干净数据和含异常值数据下均优于当前最优的UFS方法。

原文摘要 · Abstract (English)

Effective feature selection is essential for high-dimensional data analysis and machine learning. Unsupervised feature selection (UFS) aims to simultaneously cluster data and identify the most discriminative features. Most existing UFS methods linearly project features into a pseudo-label space for clustering, but they suffer from two critical limitations: (1) an oversimplified linear mapping that fails to capture complex feature relationships, and (2) an assumption of uniform cluster distributions, ignoring outliers prevalent in real-world data. To address these issues, we propose the Robust Autoencoder-based Unsupervised Feature Selection (RAEUFS) model, which leverages a deep autoencoder to learn nonlinear feature representations while inherently improving robustness to outliers. We further develop an efficient optimization algorithm for RAEUFS. Extensive experiments demonstrate that our method outperforms state-of-the-art UFS approaches in both clean and outlier-contaminated data settings.

特征选择自编码器无监督鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。