arXiv:2602.05749cs.LG2026-02被引 1

不用深度学习也能实现深度聚类的目标,只需利用数据分布信息。

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

  • 用数据分布信息替代深度网络来改进聚类
  • 非深度方法在形状/密度/大小各异的聚类上表现更优
  • 适合关注聚类本质机制的研究者

深度聚类常被宣称优于k-means聚类,但这种优势多基于图像数据集,且未解决k-means的根本缺陷:无法发现任意形状、不同密度和大小的聚类。深度嵌入聚类(DEC)通过自编码器学习潜在表示,并以类似k-means的方式进行聚类,优化过程端到端完成。本文研究发现,尽管采用深度学习,DEC仍未能克服k-means的这些根本局限。更重要的是,现有深度聚类方法均未有效利用数据内在分布。本研究揭示:仅依靠数据分布信息的非深度方法,即可实现深度聚类的预期目标,有效应对上述挑战。该发现对深度聚类方法具有广泛启示。

原文摘要 · Abstract (English)

Deep clustering (DC) is often quoted to have a key advantage over $k$-means clustering. Yet, this advantage is often demonstrated using image datasets only, and it is unclear whether it addresses the fundamental limitations of $k$-means clustering. Deep Embedded Clustering (DEC) learns a latent representation via an autoencoder and performs clustering based on a $k$-means-like procedure, while the optimization is conducted in an end-to-end manner. This paper investigates whether the deep-learned representation has enabled DEC to overcome the known fundamental limitations of $k$-means clustering, i.e., its inability to discover clusters of arbitrary shapes, varied sizes and densities. Our investigations on DEC have a wider implication on deep clustering methods in general. Notably, none of these methods exploit the underlying data distribution. We uncover that a non-deep learning approach achieves the intended aim of deep clustering by making use of distributional information of clusters in a dataset to effectively address these fundamental limitations.

聚类深度学习数据分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。