提出新模型FSbuHD,用距离优化模糊粗糙集特征选择。
A New Modeling to Feature Selection Based on the Fuzzy Rough Set Theory in Normal and Optimistic States on Hybrid Information Systems
- 基于对象间距离构建模糊等价关系,避免高维交集运算。
- 在UCI数据集上优于现有方法,正常与乐观模式均表现优异。
- 适合处理混合信息系统的高维数据特征筛选任务。
面对大数据生成的海量、多样与高速特性,研究特征选择方法具有广泛的应用价值。通过剔除无关与冗余特征,特征选择可降低数据维度,提升决策系统效率。模糊粗糙集理论是混合信息系统中特征选择的关键工具,但面临两大挑战:高维空间中通过交集运算获取模糊等价关系耗时且占用内存;该方法可能引入噪声,影响特征选择效果。本文提出一种新模型FSbuHD,通过计算对象间的联合距离来构建模糊等价关系,将特征选择问题转化为可由元启发式算法求解的优化问题。该模型支持正常与乐观两种运行模式,分别基于两种引入的模糊等价关系。在标准UCI数据集上的实验表明,与已有算法相比,FSbuHD在效率与效果上均表现出色。
原文摘要 · Abstract (English)
Considering the high volume, wide variety, and rapid speed of data generation, investigating feature selection methods for big data presents various applications and advantages. By removing irrelevant and redundant features, feature selection reduces data dimensions, thereby facilitating optimal decision-making within decision systems. One of the key tools for feature selection in hybrid information systems is fuzzy rough set theory. However, this theory faces two significant challenges: First, obtaining fuzzy equivalence relations through intersection operations in high-dimensional spaces can be both time-consuming and memory-intensive. Additionally, this method may produce noisy data, complicating the feature selection process. The purpose and innovation of this paper are to address these issues. We proposed a new feature selection model that calculates the combined distance between objects and subsequently used this information to derive the fuzzy equivalence relation. Rather than directly solving the feature selection problem, this approach reformulates it into an optimization problem that can be tackled using appropriate meta-heuristic algorithms. We have named this new approach FSbuHD. The FSbuHD model operates in two modes - normal and optimistic - based on the selection of one of the two introduced fuzzy equivalence relations. The model is then tested on standard datasets from the UCI repository and compared with other algorithms. The results of this research demonstrate that FSbuHD is one of the most efficient and effective methods for feature selection when compared to previous methods and algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。