提出量化聚类结果可复现性的新方法,解决科学分析中聚类稳定性难题。
ERICA: Quantifying Replicability of Cluster Analysis

- 通过迭代聚类分配计算可复现性统计量,判断数据集是否存在稳定聚类结构。
- 在乳腺癌基因表达数据中发现多个聚类解不可复现,揭示现有分析潜在风险。
- 适用于生物信息学、数据挖掘等领域,帮助研究人员验证聚类结果可靠性。
尽管聚类在科学中广泛应用,但缺乏统一的定量评估框架来衡量其结果的可复现性。本文提出基于迭代聚类分配的可复现性评估方法(ERICA),用于判断数据集中是否存在可重复的聚类结构。该方法通过计算统计量来检测可复现的聚类模式,并引入定量可视化手段,刻画聚类间相似性,识别可能的离群点或不稳定的分配。在合成数据集上的实验表明,ERICA能有效识别出可复现的聚类结构。而在三个乳腺癌基因表达数据集上的应用则揭示了部分聚类解不具备可复现性。研究强调了严格评估聚类结果的重要性,并提供了一个实用的评估框架。
原文摘要 · Abstract (English)
Despite being ubiquitous in science, clustering lacks a unified framework for quantitatively evaluating the replicability of its results. We present evaluating replicability via iterative clustering assignments (ERICA), a method for determining whether clusters can be identified reproducibly in a dataset. The pipeline computes a statistic that determines whether reproducible cluster structure is present in a dataset. Quantitative visualization methods are also introduced to characterize similarities between clusters and identify observations that may represent outliers or unstable assignments. Experiments on synthetic datasets demonstrate that ERICA successfully identifies reproducible cluster structure. In contrast, application of ERICA to three breast cancer gene-expression datasets reveals instances in which clustering solutions are not reproducible. The study underscores the importance of rigorously evaluating clustering solutions and provides a practical framework for doing so.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。