用视觉语言模型增强谱聚类,提升无监督图像分组效果。
Delving into Spectral Clustering with Vision-Language Representations
- 利用预训练多模态模型中的跨模态对齐,构建融合视觉与语义的相似度矩阵。
- 在16个基准上显著优于现有方法,尤其在细粒度和域偏移数据上表现突出。
- 适合需要高质量无监督图像聚类的科研与工业场景。
谱聚类是一种强大的无监督数据分析技术,但现有方法大多依赖单一模态,未能充分利用多模态表示中的丰富信息。受视觉-语言预训练成功启发,本文将谱聚类从单模态拓展至多模态范式。提出神经切线核谱聚类(Neural Tangent Kernel Spectral Clustering),通过锚定正向名词(即与目标图像语义相近的词汇)来构建视觉相似性与语义重叠耦合的图像间亲密度。该设计强化了簇内连接,抑制了跨簇的虚假关联,从而促进块对角结构。此外,提出一种正则化亲密度扩散机制,自适应融合不同提示词生成的亲密度矩阵。在16个基准数据集——涵盖经典、大规模、细粒度及域偏移数据——上的实验表明,本方法显著优于当前最优水平。
原文摘要 · Abstract (English)
Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of spectral clustering from a single-modal to a multi-modal regime. Particularly, we propose Neural Tangent Kernel Spectral Clustering that leverages cross-modal alignment in pre-trained vision-language models. By anchoring the neural tangent kernel with positive nouns, i.e., those semantically close to the images of interest, we arrive at formulating the affinity between images as a coupling of their visual proximity and semantic overlap. We show that this formulation amplifies within-cluster connections while suppressing spurious ones across clusters, hence encouraging block-diagonal structures. In addition, we present a regularized affinity diffusion mechanism that adaptively ensembles affinity matrices induced by different prompts. Extensive experiments on \textbf{16} benchmarks -- including classical, large-scale, fine-grained and domain-shifted datasets -- manifest that our method consistently outperforms the state-of-the-art by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。