arXiv:2512.00979stat.MLcs.AI2025-12

用转置数据的K均值聚类,揭示变量分组对主成分的贡献。

An Approach to Variable Clustering: K-means in Transposed Data and its Relationship with Principal Component Analysis

  • 在转置数据上做K均值,将变量当观测聚类。
  • 通过变量载荷量化各变量簇对主成分的贡献。
  • 适合想理解变量分组与主成分关系的研究者。

主成分分析(PCA)和K均值聚类是多元分析中的基础方法。尽管它们常被独立或顺序使用于观测聚类,但关于二者关系,特别是将K均值用于变量聚类而非观测聚类时的研究仍十分有限。本研究提出一种新方法:先对原始数据做PCA,再在转置数据集上应用K均值聚类(此时原变量变为观测),并通过变量载荷衡量每个变量簇对各主成分的贡献。该方法提供了一种探索变量聚类及其在主成分方向上贡献的工具,有助于理解变量分组与数据主要变异方向之间的关系。

原文摘要 · Abstract (English)

Principal Component Analysis (PCA) and K-means constitute fundamental techniques in multivariate analysis. Although they are frequently applied independently or sequentially to cluster observations, the relationship between them, especially when K-means is used to cluster variables rather than observations, has been scarcely explored. This study seeks to address this gap by proposing an innovative method that analyzes the relationship between clusters of variables obtained by applying K-means on transposed data and the principal components of PCA. Our approach involves applying PCA to the original data and K-means to the transposed data set, where the original variables are converted into observations. The contribution of each variable cluster to each principal component is then quantified using measures based on variable loadings. This process provides a tool to explore and understand the clustering of variables and how such clusters contribute to the principal dimensions of variation identified by PCA.

变量聚类主成分分析K均值数据降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。