提出新检验方法,解决聚类后子集误判问题
Two-cluster test
- 基于两子集边界点构造新统计量
- 显著降低传统方法的假阳性率
- 适用于可解释聚类与层次聚类
聚类分析是统计学与机器学习中的基础问题。在许多现代聚类方法中,需判断两个样本子集是否来自同一簇。由于这些子集通常由聚类过程生成,直接使用经典两样本检验会导致极小p值,造成显著的假阳性(Ⅰ类错误)膨胀。为此,本文首次正式定义‘双簇检验’问题,指出其与传统两样本检验本质不同。我们提出一种基于两子集边界点的新方法,可解析计算p值以量化显著性。在合成与真实数据集上的实验表明,该方法显著降低Ⅰ类错误率,优于多种经典两样本检验。更重要的是,该检验在树状可解释聚类与基于显著性的层次聚类中得到实际验证。
原文摘要 · Abstract (English)
Cluster analysis is a fundamental research issue in statistics and machine learning. In many modern clustering methods, we need to determine whether two subsets of samples come from the same cluster. Since these subsets are usually generated by certain clustering procedures, the deployment of classic two-sample tests in this context would yield extremely smaller p-values, leading to inflated Type-I error rate. To overcome this bias, we formally introduce the two-cluster test issue and argue that it is a totally different significance testing issue from conventional two-sample test. Meanwhile, we present a new method based on the boundary points between two subsets to derive an analytical p-value for the purpose of significance quantification. Experiments on both synthetic and real data sets show that the proposed test is able to significantly reduce the Type-I error rate, in comparison with several classic two-sample testing methods. More importantly, the practical usage of such two-cluster test is further verified through its applications in tree-based interpretable clustering and significance-based hierarchical clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。