统一学习图像嵌入与聚类,通过图切割引导压缩提升聚类精度。
Graph Cut-guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
- 将聚类模块融入结构化表示学习框架,实现嵌入与聚类联合优化。
- 在标准与域外数据集上均优于现有方法,提升聚类性能。
- 适合需要高质量图像聚类的场景,如无监督表征学习。
在预训练模型时代,图像聚类通常分为两个阶段:首先从预训练视觉模型提取特征,然后基于这些特征进行聚类。然而,这两个阶段常被分开处理或采用不同学习范式,导致聚类效果不理想。本文提出一种统一框架——图切割引导的最大编码速率缩减(CgMCR²),用于联合学习结构化嵌入与聚类。具体而言,将高效的聚类模块集成到结构化表示学习的原理性框架中,聚类模块提供划分信息以指导簇内压缩,而学习到的嵌入则对齐于期望的几何结构,反过来提升聚类准确性。我们在标准及域外图像数据集上进行了广泛实验,结果验证了该方法的有效性。
原文摘要 · Abstract (English)
In the era of pre-trained models, image clustering task is usually addressed by two relevant stages: a) to produce features from pre-trained vision models; and b) to find clusters from the pre-trained features. However, these two stages are often considered separately or learned by different paradigms, leading to suboptimal clustering performance. In this paper, we propose a unified framework, termed graph Cut-guided Maximal Coding Rate Reduction (CgMCR$^2$), for jointly learning the structured embeddings and the clustering. To be specific, we attempt to integrate an efficient clustering module into the principled framework for learning structured representation, in which the clustering module is used to provide partition information to guide the cluster-wise compression and the learned embeddings is aligned to desired geometric structures in turn to help for yielding more accurate partitions. We conduct extensive experiments on both standard and out-of-domain image datasets and experimental results validate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。