通过融合邻域区间扰动,提升无监督特征选择的稳定性与通用性。
Unsupervised feature selection algorithm framework based on neighborhood interval disturbance fusion
- 基于邻域区间扰动融合,实现特征评分与数据近似区间的联合学习。
- 在多个数据集上优于传统无监督方法,显著提升特征选择稳定性。
- 适合高维无标签数据处理,尤其适用于结构复杂的数据集。
特征选择是数据降维的关键技术。由于采集数据样本缺乏标签信息,无监督特征选择受到广泛关注。然而,许多现有算法通用性与稳定性较差,且受数据集结构影响较大。为此,研究者们致力于提升算法稳定性。本文通过预处理数据集并采用区间法近似数据,实验验证了新生成区间数据集的优劣。从全局视角出发,提出一种新的无监督特征选择算法——基于邻域区间扰动融合(NIDF)。该方法可实现特征最终得分与数据近似区间的联合学习。通过与原始无监督特征选择方法及若干现有框架对比,验证了所提模型的优越性。
原文摘要 · Abstract (English)
Feature selection technology is a key technology of data dimensionality reduction. Becauseof the lack of label information of collected data samples, unsupervised feature selection has attracted more attention. The universality and stability of many unsupervised feature selection algorithms are very low and greatly affected by the dataset structure. For this reason, many researchers have been keen to improve the stability of the algorithm. This paper attempts to preprocess the data set and use an interval method to approximate the data set, experimentally verifying the advantages and disadvantages of the new interval data set. This paper deals with these data sets from the global perspective and proposes a new algorithm-unsupervised feature selection algorithm based on neighborhood interval disturbance fusion(NIDF). This method can realize the joint learning of the final score of the feature and the approximate data interval. By comparing with the original unsupervised feature selection methods and several existing feature selection frameworks, the superiority of the proposed model is verified.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。