arXiv:2505.04464cs.LGcs.AI2025-05中稿 · UAI 2025

提出一种基于共识的聚类模型排序方法,能更好区分不同聚类结果优劣。

Discriminative Ordering Through Ensemble Consensus

  • 通过聚类连接性与共识矩阵的距离构建判别性排序
  • 在合成数据上准确识别最符合共识的聚类模型
  • 适用于多算法、无固定簇数且可融合约束的场景

聚类模型评估因簇定义不明确而困难,现有指标难以处理多样化的聚类定义,也难以融入约束条件。本文受共识聚类启发,假设一组聚类模型能揭示数据隐藏结构。提出基于聚类模型连接性与共识矩阵距离的集成聚类判别性排序方法。在合成数据上验证,该评分能优先排列最符合共识的模型。进一步表明,该简单排序得分在比较不同聚类算法(不限定固定簇数)时显著优于其他方法,且兼容聚类约束。

原文摘要 · Abstract (English)

Evaluating the performance of clustering models is a challenging task where the outcome depends on the definition of what constitutes a cluster. Due to this design, current existing metrics rarely handle multiple clustering models with diverse cluster definitions, nor do they comply with the integration of constraints when available. In this work, we take inspiration from consensus clustering and assume that a set of clustering models is able to uncover hidden structures in the data. We propose to construct a discriminative ordering through ensemble clustering based on the distance between the connectivity of a clustering model and the consensus matrix. We first validate the proposed method with synthetic scenarios, highlighting that the proposed score ranks the models that best match the consensus first. We then show that this simple ranking score significantly outperforms other scoring methods when comparing sets of different clustering algorithms that are not restricted to a fixed number of clusters and is compatible with clustering constraints.

聚类评估共识聚类模型排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。