用图文对比学习提升无监督图像聚类效果
Dual-Level Cross-Modal Contrastive Clustering
- 引入外部文本构建语义空间,生成图文对
- 双层次跨模态对比学习,提升特征区分度
- 在5个数据集上表现优于现有方法,适合图像聚类研究
图像聚类作为无监督学习的关键任务,旨在不依赖标签的情况下将图像分组。尽管已有深度聚类方法取得显著进展,但大多仅关注图像自身内在信息,忽略了外部语义知识对理解的增强作用。近年来,基于大规模数据预训练的视觉-语言模型在下游任务中表现优异,但视觉表征与文本语义学习之间仍存在差距。如何有效融合多模态表示用于聚类仍是挑战。为此,我们提出一种新型图像聚类框架——双层次跨模态对比聚类(DXMC)。首先,引入外部文本信息构建语义空间,并生成图像-文本配对;其次,将配对输入预训练的图像与文本编码器,获取嵌入表示,并送入四个精心设计的网络;最后,在不同层次上进行跨模态对比学习,增强判别性表示。在五个基准数据集上的大量实验表明,所提方法具有明显优势。
原文摘要 · Abstract (English)
Image clustering, which involves grouping images into different clusters without labels, is a key task in unsupervised learning. Although previous deep clustering methods have achieved remarkable results, they only explore the intrinsic information of the image itself but overlook external supervision knowledge to improve the semantic understanding of images. Recently, visual-language pre-trained model on large-scale datasets have been used in various downstream tasks and have achieved great results. However, there is a gap between visual representation learning and textual semantic learning, and how to properly utilize the representation of two different modalities for clustering is still a big challenge. To tackle the challenges, we propose a novel image clustering framwork, named Dual-level Cross-Modal Contrastive Clustering (DXMC). Firstly, external textual information is introduced for constructing a semantic space which is adopted to generate image-text pairs. Secondly, the image-text pairs are respectively sent to pre-trained image and text encoder to obtain image and text embeddings which subsquently are fed into four well-designed networks. Thirdly, dual-level cross-modal contrastive learning is conducted between discriminative representations of different modalities and distinct level. Extensive experimental results on five benchmark datasets demonstrate the superiority of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。