arXiv:2508.15444cs.LG2025-08中稿 · the 33rd Internati…

提出新方法快速检测高维高斯聚类重叠,适合流数据在线聚类。

Measures of Overlapping Multivariate Gaussian Clusters in Unsupervised Online Learning

  • 专为检测聚类重叠设计,不依赖传统相似性度量
  • 计算速度是现有方法的数倍,能准确识别重叠聚类
  • 适合处理概念漂移下的流数据聚类,避免误合并正交聚类

本文提出一种新型度量方法,用于检测多变量高斯聚类间的重叠。在线学习任务的目标是从数据流中构建可随概念漂移动态适应的聚类、分类或回归模型。在聚类场景下,数据流可能产生大量重叠的聚类,需进行合并。然而,常用分布差异度量在流数据在线学习中表现不佳,因其无法适应各类聚类形状且计算开销大。本文提出的度量专注于检测重叠而非衡量差异,计算效率显著更高。实验表明,该方法比现有方法快数倍,能有效识别重叠聚类,同时避免将正交聚类错误合并。

原文摘要 · Abstract (English)

In this paper, we propose a new measure for detecting overlap in multivariate Gaussian clusters. The aim of online learning from data streams is to create clustering, classification, or regression models that can adapt over time based on the conceptual drift of streaming data. In the case of clustering, this can result in a large number of clusters that may overlap and should be merged. Commonly used distribution dissimilarity measures are not adequate for determining overlapping clusters in the context of online learning from streaming data due to their inability to account for all shapes of clusters and their high computational demands. Our proposed dissimilarity measure is specifically designed to detect overlap rather than dissimilarity and can be computed faster compared to existing measures. Our method is several times faster than compared methods and is capable of detecting overlapping clusters while avoiding the merging of orthogonal clusters.

聚类流数据高斯模型在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。