用视觉语言模型增强图像聚类,提升类别区分力与语义可解释性
Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers
- 利用跨模态关系生成更具区分性的自监督信号
- 通过提示学习构建连续语义中心,聚类性能平均提升2.6%
- 适用于需要高可解释性聚类结果的研究场景
语言辅助图像聚类(LAIC)借助视觉语言模型(VLMs)为图像添加文本信息以提升聚类效果。现有方法常忽略两个问题:(i) 每张图像的文本特征高度相似,导致类间区分能力弱;(ii) 聚类步骤受限于预设的图文对齐,难以充分挖掘文本模态潜力。为此,我们提出新框架,包含两个互补组件:首先,利用跨模态关系生成更具区分性的自监督信号,兼容多数VLM训练机制;其次,通过提示学习学习类别级连续语义中心,完成最终聚类分配。在八个基准数据集上的实验表明,本方法相比最先进方法平均提升2.6%;所学语义中心具备强可解释性。代码见附录。
原文摘要 · Abstract (English)
Language-Assisted Image Clustering (LAIC) augments the input images with additional texts with the help of vision-language models (VLMs) to promote clustering performance. Despite recent progress, existing LAIC methods often overlook two issues: (i) textual features constructed for each image are highly similar, leading to weak inter-class discriminability; (ii) the clustering step is restricted to pre-built image-text alignments, limiting the potential for better utilization of the text modality. To address these issues, we propose a new LAIC framework with two complementary components. First, we exploit cross-modal relations to produce more discriminative self-supervision signals for clustering, as it compatible with most VLMs training mechanisms. Second, we learn category-wise continuous semantic centers via prompt learning to produce the final clustering assignments. Extensive experiments on eight benchmark datasets demonstrate that our method achieves an average improvement of 2.6% over state-of-the-art methods, and the learned semantic centers exhibit strong interpretability. Code is available in the supplementary material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。