通过正交子空间分解,提升高维数据聚类的准确性和效率。
Orthogonal Subspace Clustering: Enhancing High-Dimensional Data Analysis through Adaptive Dimensionality Reduction and Efficient Clustering
- 基于统计正交分解理论,自动选择最优降维维度。
- 在多个基准数据集上,聚类准确率、NMI等指标均优于现有方法。
- 适合处理高维稀疏数据,尤其适用于需要鲁棒聚类的场景。
本文提出正交子空间聚类(OSC),一种用于高维数据聚类的新方法。我们首先建立一个理论定理,证明高维数据可从统计意义上分解为正交子空间,其形式与Q型因子分析完全一致,为矩阵分解和因子分析驱动的降维提供了坚实的数学基础。基于该定理,我们提出OSC框架,解决因样本稀疏和距离度量失效导致的“维度灾难”问题。OSC将正交子空间构建与经典聚类方法结合,引入数据驱动机制,根据累积方差贡献自动选择子空间维度,避免人为选择偏差,同时最大化判别信息保留。通过将高维数据投影至互不相关的低维正交子空间,显著提升聚类效率、鲁棒性与准确性。在多个基准数据集上的大量实验表明,其在聚类准确率(ACC)、标准化互信息(NMI)和调整兰德指数(ARI)等指标上均优于现有方法。
原文摘要 · Abstract (English)
This paper presents Orthogonal Subspace Clustering (OSC), an innovative method for high-dimensional data clustering. We first establish a theoretical theorem proving that high-dimensional data can be decomposed into orthogonal subspaces in a statistical sense, whose form exactly matches the paradigm of Q-type factor analysis. This theorem lays a solid mathematical foundation for dimensionality reduction via matrix decomposition and factor analysis. Based on this theorem, we propose the OSC framework to address the "curse of dimensionality" -- a critical challenge that degrades clustering effectiveness due to sample sparsity and ineffective distance metrics. OSC integrates orthogonal subspace construction with classical clustering techniques, introducing a data-driven mechanism to select the subspace dimension based on cumulative variance contribution. This avoids manual selection biases while maximizing the retention of discriminative information. By projecting high-dimensional data into an uncorrelated, low-dimensional orthogonal subspace, OSC significantly improves clustering efficiency, robustness, and accuracy. Extensive experiments on various benchmark datasets demonstrate the effectiveness of OSC, with thorough analysis of evaluation metrics including Cluster Accuracy (ACC), Normalized Mutual Information (NMI), and Adjusted Rand Index (ARI) highlighting its advantages over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。