arXiv:2507.01998cs.LG2025-07

提出一种高效特征选择方法,能在个人电脑上快速处理海量数据。

Positive region preserved random sampling: an efficient feature selection method for massive data

  • 基于采样与粗糙集理论,用可区分对象对比例衡量特征集判别能力。
  • 在11个数据集上实现快速近似约简,判别能力超过预估下界。
  • 适合资源受限场景下的海量数据特征筛选,尤其适用于个人电脑环境。

特征选择是智能机器在海量数据中提升成功率的关键步骤,但计算资源常不足。本文结合采样技术与粗糙集理论,提出一种新方法:通过可区分对象对数与应区分对象对总数之比来衡量特征集的判别能力。基于此,构建保留正区域的采样样本,以快速找到具有高判别能力的特征子集。相比其他方法,该方法可在个人电脑上于可接受时间内完成特征选择,且能预先估算所选特征子集能区分的对象对概率下界。在11个不同规模的数据集上验证表明,该方法可在极短时间内找到近似约简,且最终约简的判别能力高于预估下界;在4个大规模数据集上的实验也证实,可在合理时间内获得高判别能力的近似约简。

原文摘要 · Abstract (English)

Selecting relevant features is an important and necessary step for intelligent machines to maximize their chances of success. However, intelligent machines generally have no enough computing resources when faced with huge volume of data. This paper develops a new method based on sampling techniques and rough set theory to address the challenge of feature selection for massive data. To this end, this paper proposes using the ratio of discernible object pairs to all object pairs that should be distinguished to measure the discriminatory ability of a feature set. Based on this measure, a new feature selection method is proposed. This method constructs positive region preserved samples from massive data to find a feature subset with high discriminatory ability. Compared with other methods, the proposed method has two advantages. First, it is able to select a feature subset that can preserve the discriminatory ability of all the features of the target massive data set within an acceptable time on a personal computer. Second, the lower boundary of the probability of the object pairs that can be discerned using the feature subset selected in all object pairs that should be distinguished can be estimated before finding reducts. Furthermore, 11 data sets of different sizes were used to validate the proposed method. The results show that approximate reducts can be found in a very short period of time, and the discriminatory ability of the final reduct is larger than the estimated lower boundary. Experiments on four large-scale data sets also showed that an approximate reduct with high discriminatory ability can be obtained in reasonable time on a personal computer.

特征选择粗糙集采样方法海量数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。