提出多尺度聚类显著性检验方法,发现隐藏在不同分辨率下的真实结构。
The elbow statistic: Multiscale clustering statistical significance
- 基于异质性序列构建归一化曲率统计量,形式化肘部法则
- 在合成与真实数据上控制假阳性,检测出单分辨率方法遗漏的多尺度结构
- 适用于硬聚类、模糊聚类及模型基方法,兼容性强
聚类数量选择仍是无监督学习中的核心挑战。现有方法多聚焦于寻找单一‘最优’划分,常忽略跨多分辨率的统计显著结构。本文提出ElbowSig,一种通用的推断框架,用于评估多分辨率下的聚类结构。该方法通过基于簇内异质性序列的归一化离散曲率统计量,形式化肘部法则,并相对于无结构数据的零分布评估其显著性,从而实现多尺度同时推断。我们推导了该零统计量在大样本和高维情形下的渐近行为,刻画其极限形式与变异性。由于仅依赖异质性序列,ElbowSig兼容多种聚类算法,包括硬聚类、模糊聚类及模型基方法。在合成与真实数据上的实验表明,该方法在无结构数据下能有效控制第一类错误,同时具备检测多尺度组织结构的能力,揭示出单分辨率选择标准常遗漏的结构。
原文摘要 · Abstract (English)
Selecting the number of clusters remains a fundamental challenge in unsupervised learning. Existing approaches typically focus on identifying a single "optimal" partition, often overlooking statistically meaningful structure present across multiple resolutions. We introduce ElbowSig, a general inferential framework for assessing clustering structure over a range of resolutions. The method formalizes the elbow heuristic by defining a normalized discrete curvature statistic based on the sequence of within-cluster heterogeneity values, and evaluates its significance relative to a null distribution of unstructured data. This yields hypothesis tests across resolutions, enabling simultaneous inference at multiple clustering scales. We derive the asymptotic behavior of the null statistic in both large-sample and high-dimensional regimes, characterizing its limiting form and variability. Because it depends only on the heterogeneity sequence, ElbowSig is compatible with a wide range of clustering algorithms, including hard, fuzzy, and model-based methods. Experiments on synthetic and real datasets show that the procedure controls Type-I error under unstructured data while providing power to detect multiscale organization, revealing structure that is often missed by single-resolution selection criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。