为每条数据分配独立异质参数,提升真实数据聚类效果。
Individual-heterogeneous sub-Gaussian Mixture Models
- 为每个样本引入独立异质性参数,突破传统模型同质假设。
- 在高维数据下仍能精确恢复真实聚类标签,条件宽松。
- 适用于异质性强的现实数据,尤其适合高维场景。
经典高斯混合模型假设簇内同质,但在真实数据中,观测值常表现出不同尺度或强度。为此,我们提出个体异质亚高斯混合模型,为每个观测分配独立的异质性参数,显式捕捉实际应用中的异质性。基于该模型,我们设计了一种高效谱方法,在温和分离条件下可证明实现真聚类标签的精确恢复,即使在特征数远超样本数的高维情形下依然有效。在合成与真实数据上的数值实验表明,该方法始终优于现有聚类算法,包括针对经典高斯混合模型设计的方法。
原文摘要 · Abstract (English)
The classical Gaussian mixture model assumes homogeneity within clusters, an assumption that often fails in real-world data where observations naturally exhibit varying scales or intensities. To address this, we introduce the individual-heterogeneous sub-Gaussian mixture model, a flexible framework that assigns each observation its own heterogeneity parameter, thereby explicitly capturing the heterogeneity inherent in practical applications. Built upon this model, we propose an efficient spectral method that provably achieves exact recovery of the true cluster labels under mild separation conditions, even in high-dimensional settings where the number of features far exceeds the number of samples. Numerical experiments on both synthetic and real data demonstrate that our method consistently outperforms existing clustering algorithms, including those designed for classical Gaussian mixture models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。