K-means聚类可能误把连续心理空间当分组,需谨慎解读。
Drawing Lines in Psychological Space: What K-means Clustering Reveals in Simulated and Real Psychometric Data
- 用模拟数据检验K-means在无真实分组时仍能生成稳定聚类
- 在35国大学生问卷数据中发现几何聚类模式与潜在分组不一致
- 提醒研究者警惕聚类结果的伪结构,尤其在心理特质连续分布时
K-means聚类广泛用于心理学和心理测量学中识别类型、子群和潜在分类,但其经典形式并不检验这些群体是否为潜在的心理类别。相反,K-means将多维空间划分为围绕质心的区域,偏好紧凑、近似球形的簇,由几何距离定义。本文通过一系列受控的模拟数据集检验这一局限性,并进一步扩展到包含来自35个国家大学生调查回应的SMARVUS大规模国际心理测量数据集,评估经验心理数据中是否出现类似的几何划分模式。通过对比模拟与实证数据,本文认为,即使在没有真实子群结构的连续高斯潜空间中,K-means仍可能产生稳定且视觉上连贯的聚类解。
原文摘要 · Abstract (English)
K-means clustering is widely used in psychological and psychometric research to identify profiles, subgroups, and potential typologies, yet its classical formulation does not test whether such groups exist as latent psychological categories. Instead, K-means partitions multidimensional space into regions around centroids, favoring compact, approximately spherical clusters defined by geometric distance. In this paper, we examine this limitation through a sequence of controlled simulated datasets. We then extend the analysis to the SMARVUS dataset, a large international psychometric dataset comprising survey responses from university students across 35 countries, to evaluate whether similar geometric partitioning patterns emerge in empirical psychological data. By contrasting simulated and empirical data, this paper argues that K-means can produce stable and visually coherent clustering solutions even in continuous Gaussian latent spaces without true subgroup structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。