arXiv:2505.18918stat.MLcs.LG2025-05中稿 · ance被引 1

针对噪声质量不同的数据,提出新聚类方法提升降维精度。

ALPCAHUS: Subspace Clustering for Heteroscedastic Data

  • 基于异方差性建模,估计每条数据的噪声方差。
  • 在真实与模拟数据上,聚类准确率优于传统方法。
  • 适合处理噪声差异大的多子空间数据,如传感器或图像数据。

主成分分析(PCA)是数据降维的重要工具。为应对来自多个子空间的数据聚类问题,已有方法如K-Subspaces(KSS)被提出。然而,某些应用中的数据因样本间噪声特性不同而呈现异方差性。本文提出一种基于异方差性的子空间聚类方法ALPCAHUS,可估计样本级噪声方差,并利用该信息优化数据低秩结构对应的子空间基估计。该算法在KSS框架基础上,扩展了近期提出的异方差性PCA方法LR-ALPCAH,适用于存在异方差噪声的子空间并集(UoS)场景。仿真与真实数据实验表明,考虑异方差性显著提升了聚类性能。代码已开源:https://github.com/javiersc1/ALPCAHUS。

原文摘要 · Abstract (English)

Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend PCA to the union of subspace (UoS) setting for clustering data that comes from multiple subspaces like K-Subspaces (KSS). However, some applications involve heterogeneous data that vary in quality due to noise characteristics associated with each data sample. Heteroscedastic methods aim to deal with such mixed data quality. This paper develops a heteroscedastic-based subspace clustering method, named ALPCAHUS, that can estimate the sample-wise noise variances and use this information to improve the estimate of the subspace bases associated with the low-rank structure of the data. This clustering algorithm builds on K-Subspaces (KSS) principles by extending the recently proposed heteroscedastic PCA method, named LR-ALPCAH, for clusters with heteroscedastic noise in the UoS setting. Simulations and real-data experiments show the effectiveness of accounting for data heteroscedasticity compared to existing clustering algorithms. Code available at https://github.com/javiersc1/ALPCAHUS.

子空间聚类异方差性降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。