arXiv:2511.05983stat.MLcs.LG2025-11被引 1

全面评测26种聚类有效性指标,提升评估可靠性

Benchmarking of Clustering Validity Measures Revisited

  • 设计三套互补评估方法,避免单一测试偏差
  • 使用1.6万+数据集与8种算法,覆盖多样聚类场景
  • 新数据集与方法可复用,适合评估聚类算法性能

聚类有效性评估在聚类过程中至关重要。现有多种内部有效性指标可用于从不同算法或超参数产生的候选解中选出最优聚类结果。本文对26种内部有效性指标进行了全面基准测试,涵盖经典与近年提出的方法。研究基于Vendramin等(2010)的方法进行改进,提出三种定制化互补评估子方法,分别聚焦不同行为特征,同时避免彼此干扰。每种方法包含两个性能度量,并支持深入分析指标复杂行为。此外,构建了包含16177个数据集的新数据集集合,搭配八种广泛使用的聚类算法,以增强适用性与多样性,覆盖更广泛的聚类应用场景。

原文摘要 · Abstract (English)

Validation plays a crucial role in the clustering process. Many different internal validity indexes exist for the purpose of determining the best clustering solution(s) from a given collection of candidates, e.g., as produced by different algorithms or different algorithm hyper-parameters. In this study, we present a comprehensive benchmark study of 26 internal validity indexes, which includes highly popular classic indexes as well as more recently developed ones. We adopted an enhanced revision of the methodology presented in Vendramin et al. (2010), developed here to address several shortcomings of this previous work. This overall new approach consists of three complementary custom-tailored evaluation sub-methodologies, each of which has been designed to assess specific aspects of an index's behaviour while preventing potential biases of the other sub-methodologies. Each sub-methodology features two complementary measures of performance, alongside mechanisms that allow for an in-depth investigation of more complex behaviours of the internal validity indexes under study. Additionally, a new collection of 16177 datasets has been produced, paired with eight widely-used clustering algorithms, for a wider applicability scope and representation of more diverse clustering scenarios.

聚类评估有效性指标基准测试数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。