arXiv:2501.15542cs.LG2025-01被引 178

用轮廓系数优化类别数据聚类的最优簇数

Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient

  • 基于核密度估计定义类别数据聚类中心
  • 信息论距离衡量对象与中心差异,轮廓系数选最优k
  • 在合成与真实数据集上优于其他算法

确定聚类数量(k)是划分聚类中的主要挑战之一。本文提出一种名为k-SCC的算法,用于估计类别数据聚类的最优k。该算法采用核密度估计方法定义聚类中心,并使用基于信息论的相异度度量计算中心与各簇中对象之间的距离。随后,通过轮廓系数分析评估不同聚类结果的质量,以选择最优的k。在合成数据集和真实数据集上的对比实验表明,k-SCC在确定每组数据的最优聚类数方面优于三种对比算法。

原文摘要 · Abstract (English)

The problem of estimating the number of clusters (say k) is one of the major challenges for the partitional clustering. This paper proposes an algorithm named k-SCC to estimate the optimal k in categorical data clustering. For the clustering step, the algorithm uses the kernel density estimation approach to define cluster centers. In addition, it uses an information-theoretic based dissimilarity to measure the distance between centers and objects in each cluster. The silhouette analysis based approach is then used to evaluate the quality of different clustering obtained in the former step to choose the best k. Comparative experiments were conducted on both synthetic and real datasets to compare the performance of k-SCC with three other algorithms. Experimental results show that k-SCC outperforms the compared algorithms in determining the number of clusters for each dataset.

聚类分析轮廓系数类别数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。