arXiv:2410.06265cs.LG2024-10中稿 · ICDM 2024被引 4

提出首个将密度连通性融入损失函数的深度聚类方法,自动识别复杂形状簇。

SHADE: Deep Density-based Clustering

  • 通过密度连通性损失函数学习增强分离性的表示
  • 在非高斯簇数据上聚类质量显著优于现有方法
  • 无需人工干预,可自动识别噪声点并保留簇形用于可视化

在高维噪声数据中检测任意形状的聚类对现有聚类方法构成挑战。我们提出 SHADE(Structure-preserving High-dimensional Analysis with Density-based Exploration),首个将密度连通性纳入损失函数的深度聚类算法。与现有深度聚类方法类似,SHADE 利用深度自编码器处理高维大规模数据。不同于依赖中心点目标的多数方法,SHADE 引入新型损失函数以捕捉密度连通性,从而学习能增强密度连通聚类分离性的表示。SHADE 可完全自动地检测稳定聚类和噪声点,无需用户输入。在包含非高斯簇(如视频数据)的数据上,其聚类质量显著优于现有方法。此外,SHADE 的嵌入空间保留了簇的个体形状,适合聚类结果的可视化与解释。

原文摘要 · Abstract (English)

Detecting arbitrarily shaped clusters in high-dimensional noisy data is challenging for current clustering methods. We introduce SHADE (Structure-preserving High-dimensional Analysis with Density-based Exploration), the first deep clustering algorithm that incorporates density-connectivity into its loss function. Similar to existing deep clustering algorithms, SHADE supports high-dimensional and large data sets with the expressive power of a deep autoencoder. In contrast to most existing deep clustering methods that rely on a centroid-based clustering objective, SHADE incorporates a novel loss function that captures density-connectivity. SHADE thereby learns a representation that enhances the separation of density-connected clusters. SHADE detects a stable clustering and noise points fully automatically without any user input. It outperforms existing methods in clustering quality, especially on data that contain non-Gaussian clusters, such as video data. Moreover, the embedded space of SHADE is suitable for visualization and interpretation of the clustering results as the individual shapes of the clusters are preserved.

深度聚类密度聚类无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。