量化谱聚类在数据扰动和缺失下的不确定性,提升聚类可靠性。
Quantifying uncertainty in spectral clusterings: expectations for perturbed and incomplete data
- 基于随机集理论建模数据扰动与缺失,用蒙特卡洛模拟期望聚类结果。
- 提出可计算的不确定性度量,在无限数据和样本下具一致性。
- 适合处理实验数据噪声或缺失场景的科研人员参考。
谱聚类是一种流行的无监督学习方法,能将无标签数据划分为形状各异的互斥簇。然而,实际数据常来自实验测量,存在误差或部分缺失,这些不确定性会传递至聚类结果,导致聚类不可靠。本文将不确定性建模为随机过程,基于随机集理论,提出一种计算蒙特卡洛近似统计期望聚类的数学框架,适用于被扰动、不完整甚至冗余的数据。我们提出了若干可计算的感兴趣量,并分析其在数据点趋于无穷和蒙特卡洛样本趋于无穷时的一致性。数值实验展示了所提量的性能并进行对比。
原文摘要 · Abstract (English)
Spectral clustering is a popular unsupervised learning technique which is able to partition unlabelled data into disjoint clusters of distinct shapes. However, the data under consideration are often experimental data, implying that the data is subject to measurement errors and measurements may even be lost or invalid. These uncertainties in the corrupted input data induce corresponding uncertainties in the resulting clusters, and the clusterings thus become unreliable. Modelling the uncertainties as random processes, we discuss a mathematical framework based on random set theory for the computational Monte Carlo approximation of statistically expected clusterings in case of corrupted, i.e., perturbed, incomplete, and possibly even additional, data. We propose several computationally accessible quantities of interest and analyze their consistency in the infinite data point and infinite Monte Carlo sample limit. Numerical experiments are provided to illustrate and compare the proposed quantities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。