arXiv:2601.11160cs.LGcs.AI2026-01综述

平衡抽象与表征,提升高维数据聚类效果

Clustering High-dimensional Data: Balancing Abstraction and Representation Tutorial at AAAI 2026

  • 通过潜空间分离关键聚类信息与冗余信息
  • 用中心点和密度损失显式约束抽象性
  • 适合研究聚类算法设计与高维数据处理者

如何从大规模真实数据集中发现自然分组?聚类需在抽象与表征间取得平衡:既要忽略个体细节,又要保留群体间的关键特征。传统K-means以高抽象性(平均掉细节)和简单表征(原始空间中的高斯分布)实现聚类。针对高维复杂数据,子空间聚类与深度聚类方法通过学习更丰富的表征来提升性能。但表征能力增强后,必须在目标函数中显式引入抽象约束,否则易退化为表示学习。当前深度聚类通过基于中心点和密度的损失函数来强制抽象。子空间聚类思想进一步将数据分解为聚类相关与非相关两个潜空间。未来方向是自适应地动态平衡抽象与表征,以提升性能、能效与可解释性。

原文摘要 · Abstract (English)

How to find a natural grouping of a large real data set? Clustering requires a balance between abstraction and representation. To identify clusters, we need to abstract from superfluous details of individual objects. But we also need a rich representation that emphasizes the key features shared by groups of objects that distinguish them from other groups of objects. Each clustering algorithm implements a different trade-off between abstraction and representation. Classical K-means implements a high level of abstraction - details are simply averaged out - combined with a very simple representation - all clusters are Gaussians in the original data space. We will see how approaches to subspace and deep clustering support high-dimensional and complex data by allowing richer representations. However, with increasing representational expressiveness comes the need to explicitly enforce abstraction in the objective function to ensure that the resulting method performs clustering and not just representation learning. We will see how current deep clustering methods define and enforce abstraction through centroid-based and density-based clustering losses. Balancing the conflicting goals of abstraction and representation is challenging. Ideas from subspace clustering help by learning one latent space for the information that is relevant to clustering and another latent space to capture all other information in the data. The tutorial ends with an outlook on future research in clustering. Future methods will more adaptively balance abstraction and representation to improve performance, energy efficiency and interpretability. By automatically finding the sweet spot between abstraction and representation, the human brain is very good at clustering and other related tasks such as single-shot learning. So, there is still much room for improvement.

聚类深度学习表征学习潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。