用同义语义空间提升视觉语言模型零样本泛化能力
$S^3$: Synonymous Semantic Space for Improving Zero-Shot Generalization of Vision-Language Models
- 为每类图像构建由同义概念组成的连续语义空间
- 在17个基准上显著优于现有方法,最高提升6.2%
- 适合需要强零样本泛化能力的视觉任务研究者
近期研究致力于通过缓解下游任务中图像与文本嵌入的语义错位来提升视觉语言模型(如CLIP)的零样本泛化能力。然而,现有方法很少考虑自然语言中常见的词汇多样性——同一类图像可用显著不同的文本概念描述,这严重影响了CLIP的零样本性能。为此,本文提出为每个图像类别构建一个同义语义空间($S^3$),而非依赖单一文本概念,从而实现更稳定的语义对齐并提升零样本泛化能力。具体而言,$S^3$首先利用大语言模型基于类别标签生成多个同义概念,并基于这些概念的Vietoris-Rips复形构建连续而紧凑的同义语义空间;进一步探索多种点到空间度量,提出点到局部中心度量以计算图像嵌入与类别同义语义空间间的相似性,实现高效的零样本预测。在17个基准上的广泛实验表明,包括细粒度零样本分类、自然分布零样本分类和开放词汇分割,$S^3$均优于现有最优方法。
原文摘要 · Abstract (English)
Recently, many studies have been conducted to enhance the zero-shot generalization ability of vision-language models (e.g., CLIP) by addressing the semantic misalignment between image and text embeddings in downstream tasks. Although many efforts have been made, existing methods barely consider the fact that a class of images can be described by notably different textual concepts due to well-known lexical variation in natural language processing, which heavily affects the zero-shot generalization of CLIP. Therefore, this paper proposes a \textbf{S}ynonymous \textbf{S}emantic \textbf{S}pace ($S^3$) for each image class, rather than relying on a single textual concept, achieving more stable semantic alignment and improving the zero-shot generalization of CLIP. Specifically, our $S^3$ method first generates several synonymous concepts based on the label of each class by using large language models, and constructs a continuous yet compact synonymous semantic space based on the Vietoris-Rips complex of the generated synonymous concepts. Furthermore, we explore the effect of several point-to-space metrics on our $S^3$, while presenting a point-to-local-center metric to compute similarity between image embeddings and the synonymous semantic space of each class, accomplishing effective zero-shot predictions. Extensive experiments are conducted across 17 benchmarks, including fine-grained zero-shot classification, natural distribution zero-shot classification, and open-vocabulary segmentation, and the results show that our $S^3$ outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。