将图像物体与文本描述的层次关系融入超球空间表示学习,提升模型泛化能力。
Compositional Entailment Learning for Hyperbolic Vision-Language Models
- 通过图像框与文本描述的组合结构,构建层次化嵌入
- 在百万级数据上实现优于欧氏与现有超球方法的零样本性能
- 适合关注视觉语言模型层次建模的研究者
图像-文本表征学习是视觉语言模型的核心,通常将图像与文本描述在共享嵌入空间中进行对比对齐。由于视觉和文本概念天然具有层次性,近期研究证明超球空间可作为高潜力流形,用于学习具有强下游性能的视觉语言表征。本文首次展示如何通过超越单一图像-文本对的方式,充分挖掘超球嵌入的固有层次特性。提出一种面向超球视觉语言模型的组合蕴含学习方法:一张图像不仅由一句文本描述,更由多个带有各自文本描述的物体框构成。这些信息可通过提取句子中的名词并结合公开的局部定位模型免费获取。我们通过对比和蕴含目标,实现了图像、图像框及其文本描述的层次组织。在使用数百万图像-文本对训练的超球视觉语言模型上进行实证评估,结果表明所提组合学习方法优于传统的欧氏CLIP学习以及近期的超球替代方案,在零样本和检索泛化方面表现更优,且层次性能显著提升。
原文摘要 · Abstract (English)
Image-text representation learning forms a cornerstone in vision-language models, where pairs of images and textual descriptions are contrastively aligned in a shared embedding space. Since visual and textual concepts are naturally hierarchical, recent work has shown that hyperbolic space can serve as a high-potential manifold to learn vision-language representation with strong downstream performance. In this work, for the first time we show how to fully leverage the innate hierarchical nature of hyperbolic embeddings by looking beyond individual image-text pairs. We propose Compositional Entailment Learning for hyperbolic vision-language models. The idea is that an image is not only described by a sentence but is itself a composition of multiple object boxes, each with their own textual description. Such information can be obtained freely by extracting nouns from sentences and using openly available localized grounding models. We show how to hierarchically organize images, image boxes, and their textual descriptions through contrastive and entailment-based objectives. Empirical evaluation on a hyperbolic vision-language model trained with millions of image-text pairs shows that the proposed compositional learning approach outperforms conventional Euclidean CLIP learning, as well as recent hyperbolic alternatives, with better zero-shot and retrieval generalization and clearly stronger hierarchical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。