提出绝对聚类指标,直接判断聚类的紧凑性与分离度。
Absolute indices for determining compactness, separability and number of clusters
- 基于每个簇的紧凑函数和簇间邻点集,定义新指标。
- 在多种数据集上验证,能准确识别真实聚类数。
- 适合需要自动确定聚类数的场景,如无监督学习。
发现数据集中“真实”聚类是一个挑战。不同模型和算法得到的聚类结果未必具有高紧凑性和良好分离性,也未必是最佳聚类数。常用的聚类有效性指标多为相对指标,用于比较算法或调参,且其效果依赖于数据结构。本文提出新型绝对聚类指标,用于判断聚类的紧凑性与分离性。为每个簇定义紧凑函数,为簇对定义邻点集,用于衡量单个簇的紧凑性及整体分布;邻点集还用于定义簇间间隔与整体分布间隔。所提紧凑性与分离性指标可用于识别真实聚类数。在多个合成与真实数据集上测试并对比主流指标,结果表明其性能更优。
原文摘要 · Abstract (English)
Finding "true" clusters in a data set is a challenging problem. Clustering solutions obtained using different models and algorithms do not necessarily provide compact and well-separated clusters or the optimal number of clusters. Cluster validity indices are commonly applied to identify such clusters. Nevertheless, these indices are typically relative, and they are used to compare clustering algorithms or choose the parameters of a clustering algorithm. Moreover, the success of these indices depends on the underlying data structure. This paper introduces novel absolute cluster indices to determine both the compactness and separability of clusters. We define a compactness function for each cluster and a set of neighboring points for cluster pairs. This function is utilized to determine the compactness of each cluster and the whole cluster distribution. The set of neighboring points is used to define the margin between clusters and the overall distribution margin. The proposed compactness and separability indices are applied to identify the true number of clusters. Using a number of synthetic and real-world data sets, we demonstrate the performance of these new indices and compare them with other widely-used cluster validity indices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。