解析判别式聚类的发展脉络与信息论的作用
A Tutorial on Discriminative Clustering and Mutual Information
- 从决策边界到不变性判别,聚焦判别式聚类的假设演进
- 互信息是推动深度判别聚类发展的核心机制
- 适合对聚类理论与方法演进感兴趣的科研人员
聚类旨在将样本划分为具有内在一致性的组。当前聚类算法的差异主要源于对'一致性'的不同假设,可分为生成式与判别式两类。过去十年中,深度聚类方法借助神经网络处理高维数据,大多采用判别式假设。本文旨在提供判别式聚类方法的历史发展视角,重点分析其假设如何从决策边界演变为不变性判别。特别强调互信息在判别式聚类(尤其是深度聚类)中的核心作用。同时指出互信息的已知局限,并探讨判别式聚类如何应对。最后讨论聚类数选择的挑战,并通过我们开发的 Python 工具包 GemClus 展示相关技术。
原文摘要 · Abstract (English)
To cluster data is to separate samples into distinctive groups that should ideally have some cohesive properties. Today, numerous clustering algorithms exist, and their differences lie essentially in what can be perceived as ``cohesive properties''. Therefore, hypotheses on the nature of clusters must be set: they can be either generative or discriminative. As the last decade witnessed the impressive growth of deep clustering methods that involve neural networks to handle high-dimensional data often in a discriminative manner; we concentrate mainly on the discriminative hypotheses. In this paper, our aim is to provide an accessible historical perspective on the evolution of discriminative clustering methods and notably how the nature of assumptions of the discriminative models changed over time: from decision boundaries to invariance critics. We notably highlight how mutual information has been a historical cornerstone of the progress of (deep) discriminative clustering methods. We also show some known limitations of mutual information and how discriminative clustering methods tried to circumvent those. We then discuss the challenges that discriminative clustering faces with respect to the selection of the number of clusters. Finally, we showcase these techniques using the dedicated Python package, GemClus, that we have developed for discriminative clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。