arXiv:2412.00440cs.CV2024-12CVPR被引 14

让视觉语言模型摆脱单一文本束缚,实现多视角图文对齐

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

  • 用图像生成多角度文本,打破单一对齐局限
  • 多分支编码+多对多对比学习,提升图文匹配精度
  • 适合需要丰富语义理解的视觉任务研究者

在快速发展的视觉语言模型领域,对比语言图像预训练(CLIP)已取得显著进展,成为众多下游任务的基础。然而,依赖于一对一(图像,文本)对比范式从大规模杂乱网络数据中学习对齐,使CLIP面临严重的短视困境,导致对单调短文本的偏好和浅层视觉表达。为克服这些问题,本文将CLIP推进至一种全新的整体性范式,通过更新多样数据与对齐优化策略。为低成本获取丰富数据,我们采用图像到文本的自动标注方法,从多个视角、粒度和层级为每张图像生成多文本。提出两个机制以促进文本多样性。为匹配此类(图像,多文本)对,我们将CLIP图像编码器改为多分支结构,并提出多对多对比优化方法,实现图像-文本部分间的细粒度匹配。结果表明,每个图像均能学习到多样化的视觉嵌入,带来良好的可解释性与泛化能力。在超过十个基准上的大量实验与消融分析显示,所提出的整体性CLIP显著优于现有短视型CLIP,涵盖图像-文本检索、开放词汇分类及密集视觉任务。

原文摘要 · Abstract (English)

In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web data, CLIP faces a serious myopic dilemma, resulting in biases towards monotonous short texts and shallow visual expressivity. To overcome these issues, this paper advances CLIP into one novel holistic paradigm, by updating both diverse data and alignment optimization. To obtain colorful data with low cost, we use image-to-text captioning to generate multi-texts for each image, from multiple perspectives, granularities, and hierarchies. Two gadgets are proposed to encourage textual diversity. To match such (image, multi-texts) pairs, we modify the CLIP image encoder into multi-branch, and propose multi-to-multi contrastive optimization for image-text part-to-part matching. As a result, diverse visual embeddings are learned for each image, bringing good interpretability and generalization. Extensive experiments and ablations across over ten benchmarks indicate that our holistic CLIP significantly outperforms existing myopic CLIP, including image-text retrieval, open-vocabulary classification, and dense visual tasks.

视觉语言模型对比学习多文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。