arXiv:2604.01943stat.MLcs.LG2026-04

提出新方法在未知聚类数下准确聚类异方差高斯数据。

A Novel Theoretical Analysis for Clustering Heteroscedastic Gaussian Data without Knowledge of the Number of Clusters

  • 基于沃尔德检验构建新代价函数,通过固定点估计聚类中心。
  • 理论证明当样本量和中心间距足够大时,固定点收敛于真实中心。
  • 新算法CENTRE-X复杂度更低,适合高维数据且无需预知聚类数。

本文研究异方差高斯数据的聚类问题,即同一簇内测量向量具有不同且未知的协方差矩阵。假设簇内数据服从以簇中心为均值的高斯分布,提出一种新的代价函数来估计中心位置。该代价函数梯度为零的点即为某函数的不动点,方法推广了已有均值漂移(Mean-Shift)算法的思路。主要理论贡献在于:当每簇样本量足够多且簇间距离足够远时,该函数的唯一不动点趋于真实聚类中心。其次,引入沃尔德核(Wald kernel),定义为对高斯均值进行沃尔德检验的p值,衡量数据属于某簇的合理性,其在高维下表现优于传统高斯核。最终,基于该理论框架提出新算法CENTRE-X,通过求解不动点实现聚类,同样无需预先知道聚类数量。与均值漂移相比,利用沃尔德检验显著减少需计算的不动点数量,降低计算复杂度。合成与真实数据集上的实验表明,即使协方差未知,CENTRE-X性能也优于或相当于K-means和均值漂移。

原文摘要 · Abstract (English)

This paper addresses the problem of clustering measurement vectors that are heteroscedastic in that they can have different covariance matrices. From the assumption that the measurement vectors within a given cluster are Gaussian distributed with possibly different and unknown covariant matrices around the cluster centroid, we introduce a novel cost function to estimate the centroids. The zeros of the gradient of this cost function turn out to be the fixed-points of a certain function. As such, the approach generalizes the methodology employed to derive the existing Mean-Shift algorithm. But as a main and novel theoretical result compared to Mean-Shift, this paper shows that the sole fixed-points of the identified function tend to be the cluster centroids if both the number of measurements per cluster and the distances between centroids are large enough. As a second contribution, this paper introduces the Wald kernel for clustering. This kernel is defined as the p-value of the Wald hypothesis test for testing the mean of a Gaussian. As such, the Wald kernel measures the plausibility that a measurement vector belongs to a given cluster and it scales better with the dimension of the measurement vectors than the usual Gaussian kernel. Finally, the proposed theoretical framework allows us to derive a new clustering algorithm called CENTRE-X that works by estimating the fixed-points of the identified function. As Mean-Shift, CENTRE-X requires no prior knowledge of the number of clusters. It relies on a Wald hypothesis test to significantly reduce the number of fixed points to calculate compared to the Mean-Shift algorithm, thus resulting in a clear gain in complexity. Simulation results on synthetic and real data sets show that CENTRE-X has comparable or better performance than standard clustering algorithms K-means and Mean-Shift, even when the covariance matrices are not perfectly known.

聚类高斯模型无监督学习统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。