用熵最小化追踪数据流中的聚类与异常点
Incremental Gaussian Mixture Clustering for Data Streams
- 基于熵最小化动态更新聚类结构
- 可同时发现聚类与远离集群的异常点
- 适用于大规模实时数据流分析
大规模数据流的分析问题在多个应用领域中至关重要。本文提出并验证了一种算法,用于在流式数据中发现聚类和异常数据点。该方法以熵最小化为准则,定义并更新由流式数据形成的聚类。随着聚类的形成,算法还能识别出远离所有已知聚类的数据点作为异常。通过多个二维数据集的实验,验证了该方法在发现聚类及识别异常点方面的有效性。
原文摘要 · Abstract (English)
The problem of analyzing data streams of very large volumes is important and is very desirable for many application domains. In this paper we present and demonstrate effective working of an algorithm to find clusters and anomalous data points in a streaming datasets. Entropy minimization is used as a criterion for defining and updating clusters formed from a streaming dataset. As the clusters are formed we also identify anomalous datapoints that show up far away from all known clusters. With a number of 2-D datasets we demonstrate the effectiveness of discovering the clusters and also identifying anomalous data points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。