系统对比五种降维方法对四种聚类算法的影响
Assessing the impact of dimensionality reduction on clustering performance -- a systematic study

- 对比PCA、VAE等五种降维法在不同维度下的表现
- 发现降维后聚类效果受数据特征和算法类型影响显著
- 适合研究降维与聚类关系的算法开发者参考
高维数据聚类前需进行降维处理,但现有研究尚未全面评估不同降维方法在多种数据类型上的效果。本研究系统比较了五种降维技术——主成分分析(PCA)、核主成分分析(Kernel PCA)、变分自编码器(VAE)、等距映射(Isomap)和多维缩放(MDS)——对四种主流聚类算法(k-means、凝聚层次聚类(AHC)、高斯混合模型(GMM)和基于点序识别聚类结构(OPTICS))的影响。采用调整兰德指数(ARI)评估聚类质量,对比无降维与在文献推荐的降维水平(即簇数减一、原始维度的25%和50%)下的结果。研究结果强调,降维方法及其维度应根据数据内在几何结构和具体聚类算法进行针对性选择。
原文摘要 · Abstract (English)
Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。