通过频域分析捕捉数据流中的概念漂移,提升变化检测精度。
Describing Nonstationary Data Streams in Frequency Domain
- 从特征向量中提取关键频率分量,按方差筛选信息量高的成分。
- 在真实数据流上实现概念识别准确率提升,优于主流方法与PCA基线。
- 适合需要实时监测数据分布变化的场景,如金融、物联网领域。
概念漂移是数据流处理中的主要挑战之一。为应对这一问题,现有漂移检测策略通常依赖于对问题元特征的分析。本文提出一种频率过滤元描述符(Frequency Filtering Metadescriptor),用于刻画数据流,通过在样本特征向量中搜索可见的有信息量的频率成分来实现。频率根据其在所有可用数据批次间的方差进行过滤。该方法可生成数据流的元描述,将数据段分组以表征特定概念,并在原始空间域中可视化频率特征。实验分析将该方案与两种前沿策略及PCA基线在事后概念识别任务中进行了对比。研究还进一步在真实数据流中识别了概念。提出的频域泛化策略能够以较少的频率成分捕捉复杂的特征依赖关系,同时保留数据的语义含义。
原文摘要 · Abstract (English)
Concept drift is among the primary challenges faced by the data stream processing methods. The drift detection strategies, designed to counteract the negative consequences of such changes, often rely on analyzing the problem metafeatures. This work presents the Frequency Filtering Metadescriptor -- a tool for characterizing the data stream that searches for the informative frequency components visible in the sample's feature vector. The frequencies are filtered according to their variance across all available data batches. The presented solution is capable of generating a metadescription of the data stream, separating chunks into groups describing specific concepts on its basis, and visualizing the frequencies in the original spatial domain. The experimental analysis compared the proposed solution with two state-of-the-art strategies and with the PCA baseline in the post-hoc concept identification task. The research is followed by the identification of concepts in the real-world data streams. The generalization in the frequency domain adapted in the proposed solution allows to capture the complex feature dependencies as a reduced number of frequency components, while maintaining the semantic meaning of data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。