针对流数据缺失值问题,提出新型进化算法实现高效稀疏特征选择。
Online Sparse Feature Selection in Data Streams via Differential Evolution
- 用潜在因子模型补全缺失数据,提升流数据处理鲁棒性。
- 基于差分进化评估特征重要性,显著提升分类准确率。
- 适合高维流数据中存在缺失值的场景,如工业传感器监测。
高维流数据处理常依赖在线流特征选择(OSFS)技术。但实际应用中常因设备故障和技术限制导致数据不完整。现有在线稀疏流特征选择(OS2FS)方法虽通过潜在因子分析进行缺失数据填补,但在特征评估方面仍存在明显不足,导致性能下降。为此,本文提出一种新的在线差分进化稀疏特征选择方法(ODESFS),包含两项关键创新:(1) 基于潜在因子分析的缺失值补全;(2) 采用差分进化算法评估特征重要性。在六个真实数据集上的全面实验表明,ODESFS在持续选择最优特征子集方面表现优异,显著优于当前主流的OSFS与OS2FS方法,整体准确率更优。
原文摘要 · Abstract (English)
The processing of high-dimensional streaming data commonly utilizes online streaming feature selection (OSFS) techniques. However, practical implementations often face challenges with data incompleteness due to equipment failures and technical constraints. Online Sparse Streaming Feature Selection (OS2FS) tackles this issue through latent factor analysis-based missing data imputation. Despite this advancement, existing OS2FS approaches exhibit substantial limitations in feature evaluation, resulting in performance deterioration. To address these shortcomings, this paper introduces a novel Online Differential Evolution for Sparse Feature Selection (ODESFS) in data streams, incorporating two key innovations: (1) missing value imputation using a latent factor analysis model, and (2) feature importance evaluation through differential evolution. Comprehensive experiments conducted on six real-world datasets demonstrate that ODESFS consistently outperforms state-of-the-art OSFS and OS2FS methods by selecting optimal feature subsets and achieving superior accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。